All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs1,992 views
Yufeng Xguard EvalA

Evaluates the safety classification capabilities of guardrail models across multiple dimensions, including prompt/response safety detection, multilingual robustness, adversarial jailbreak resilience, and safe content completion. It also tests the model's ability to dynamically adapt to new moderation policies without retraining. Use when the user wants to benchmark on Aegis / Aegis2.0, WildGuard, StrongReject, SEval2.0, E-commerce Benchmark, Adaptive Policy Scope Benchmark, or asks about eval...

researchpythongo
0
3
Yulong Me Yl MetricA

Compute yulong-me/yl_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yulong-me/yl_metric.

developmentpython
0
3
Yuyijiong Quad Match ScoreA

Compute yuyijiong/quad_match_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yuyijiong/quad_match_score.

developmentpython
0
3
Yzha Ctc EvalA

Compute yzha/ctc_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yzha/ctc_eval.

developmentpython
0
3
Zbeloki M2A

Compute zbeloki/m2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of zbeloki/m2.

developmentpython
0
3
Zebrafish Sctrans EvalA

Evaluates the ability of topological data analysis methods to capture developmental transitions and cell lineage dynamics in single-cell RNA sequencing time-series data. It specifically tests whether higher-order simplicial complexity can outperform conventional topological invariants like Betti numbers in identifying critical biological stages. Use when the user wants to benchmark on Farrell et al. (2018) zebrafish scRNA-seq, or asks about evaluating this task. Reports normalized simplicial ...

datapythonexpress
0
3
Zenbrain Memory EvalA

This evaluation protocol assesses the long-term memory and retrieval capabilities of autonomous AI systems. It measures how well models retain, route, and retrieve information across multiple sessions and varying context lengths, while also evaluating the quality of generated answers using LLM-as-a-judge scoring. Use when the user wants to benchmark on LoCoMo (Real-LoCoMo pool), LongMemEval-S, MemoryAgentBench, MemoryArena, or asks about evaluating this task. Reports NDCG@5.

ai-agentspythongo
0
3
Zero One LossA

Compute the zero_one_loss metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute zero_one_loss, or asks how to score with zero_one_loss.

documentationpythonperformance
0
3
Zero Shot Adjustable Acceleration EvalA

This evaluation protocol assesses the capability of large language models to maintain task performance while dynamically pruning hidden activations during inference. It probes the model's robustness across natural language understanding, text generation, and instruction-tuning tasks under varying computational constraints and acceleration ratios. Use when the user wants to benchmark on IMDB, GLUE, WikiText-103, Penn Treebank (PTB), One Billion Word (1BW), LAMBADA, MMLU, or asks about evaluati...

researchpythongo
0
3
Zero Shot Asr EvalA

Evaluates the impact of perceptual audio enhancement (SAM-Audio) on zero-shot automatic speech recognition performance across Bengali and English noisy speech. It probes whether signal-level quality improvements translate to better machine transcription accuracy. Use when the user wants to benchmark on Bengali Noisy YouTube dataset, English noisy dataset, or asks about evaluating this task. Reports WER, CER.

researchpythonperformance
0
3
Zero Shot Commonsense EvalA

Evaluates zero-shot commonsense reasoning capabilities of language models using multiple-choice questions. It specifically probes how prompt engineering and probability calibration strategies affect accuracy across different model sizes and architectures. Use when the user wants to benchmark on CommonsenseQA, COPA, OpenBookQA, PIQA, Social IQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Zero Shot Cot Bias EvalA

Evaluates how zero-shot Chain-of-Thought (CoT) prompting affects social bias and toxicity in large language models compared to standard prompting. It measures performance degradation on stereotype benchmarks and the propensity to generate harmful outputs on harmful question tasks. Use when the user wants to benchmark on CrowS Pairs, StereoSet, BBQ, HarmfulQ, or asks about evaluating this task. Reports TD2 accuracy.

researchpythongo
0
3
Zero Shot Generalization EvalA

Evaluates a model's ability to generalize to unseen natural language tasks without task-specific fine-tuning or prompt tuning. It probes zero-shot performance across traditional NLP benchmarks and novel BIG-bench tasks using accuracy. Use when the user wants to benchmark on BIG-bench & Held-out NLP Tasks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Zero Shot Human Classification EvalA

Evaluates the zero-shot transfer capability of a vision-language model on human-centric classification tasks, including activity recognition, age grouping, and emotion recognition, using pose-grounded text descriptions and subject-focused attention. Use when the user wants to benchmark on Stanford40, Emotic, LAGENDA-Body, LAGENDA-Face, UTKFace, FER+, or asks about evaluating this task. Reports top-k accuracy.

researchpythongo
0
3
Zero Shot Retrieval Leakage EvalA

This evaluation probes the robustness of neural retrieval models to train-test data leakage by measuring how much performance on standard benchmarks artificially improves when training data contains near-duplicates or exact matches of test queries. It specifically assesses zero-shot transfer effectiveness under varying leakage conditions and training set sizes. Use when the user wants to benchmark on Robust04, TREC 2017 Common Core, TREC 2018 Common Core, or asks about evaluating this task. R...

researchpythongo
0
3
Zero Shot Teacher Feedback EvalA

Evaluates zero-shot performance of LLMs in teacher coaching tasks, including scoring classroom transcripts against observation rubrics, identifying instructional highlights and missed opportunities, and generating actionable pedagogical suggestions. Use when the user wants to benchmark on CLASS & MQI Classroom Transcripts, or asks about evaluating this task. Reports Relevance.

researchpythongo
0
3
Zero Shot Transfer EvalA

This evaluation probes a vision-language model's ability to generalize to unseen image classification tasks without task-specific fine-tuning. It measures how well the model aligns visual features with natural language class descriptions to perform zero-shot classification across diverse domains. Use when the user wants to benchmark on ImageNet, CIFAR-10, Oxford-IIIT Pet, Food101, Stanford Cars, Kinetics700, EuroSAT, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Zero Shot Tts EvalA

Evaluates zero-shot text-to-speech synthesis capability across English and Chinese. It measures intelligibility, speaker similarity, and naturalness against reference prompts without fine-tuning on target speakers. Use when the user wants to benchmark on Seed-TTS test-en, Seed-TTS test-zh, AISHELL-3 test set, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Zero Shot Tts Vietnamese EvalA

Evaluates zero-shot text-to-speech models on Vietnamese speech generation, measuring intelligibility, speaker similarity, and naturalness across long-form and short-form text inputs. Use when the user wants to benchmark on viVoice, PAB-S, PAB-U, VIVOS, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Zero Shot Voice Synthesis EvalA

Evaluates zero-shot speech synthesis and novel voice generation by measuring how well models can produce intelligible, natural, and speaker-similar audio for unseen speakers using only conditioning embeddings. Use when the user wants to benchmark on English Multi-Accent Dataset (VCTK + Internal), or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Zeroquant EvalA

Evaluates the accuracy and inference latency of post-training quantized Transformer models (BERT and GPT-3-style) on standard NLP benchmarks and language modeling tasks. Use when the user wants to benchmark on GLUE benchmark, 20 zero-shot evaluation tasks, PTB / Wikitext-2 / Wikitext-103, or asks about evaluating this task. Reports average accuracy.

researchpythongo
0
3
Zerosense EvalA

Evaluates a model's ability to perform visual-text compression (OCR) by measuring raw text retention on rendered documents where the textual content has been deliberately stripped of semantic meaning. It isolates pure visual decoding capability from downstream linguistic priors or contextual inference. Use when the user wants to benchmark on ZeroSense, or asks about evaluating this task. Reports text preservation capability.

researchpythongo
0
3
Zest EvalA

Evaluates a model's ability to understand and generalize across unseen NLP tasks based solely on task descriptions, rather than few-shot examples. It probes systematic generalization across variations like paraphrasing, composition, semantic flips, and output structure changes. Use when the user wants to benchmark on ZEST, or asks about evaluating this task. Reports Mean.

researchpythonperformance
0
3
Zhuangbench EvalA

Evaluates large language models' ability to perform zero-shot machine translation into completely unseen, low-resource languages (Zhuang and Kalamang) using in-context learning. It probes how effectively models can adapt to new languages without prior training data by leveraging lexical expansion and syntactic exemplar retrieval. Use when the user wants to benchmark on ZhuangBench, MTOB, or asks about evaluating this task. Reports BLEU.

ai-agentspythongit
0
3
Zinc250k Mol Gen EvalA

Evaluates de novo molecular generation models on the ZINC-250k dataset across three tasks: unconditional generation from a Gaussian prior, unconstrained property optimization (maximizing penalized logP and QED), and constrained optimization (modifying molecules to improve logP while preserving structural similarity). Use when the user wants to benchmark on ZINC-250k, or asks about evaluating this task. Reports Top-3 Penalized logP.

researchpython
0
3
Zipvoice Dialog EvalA

This benchmark evaluates non-autoregressive spoken dialogue generation models on their ability to produce multi-turn conversational audio that matches input text, maintains speaker identity, and accurately handles turn-taking between two speakers. It probes both objective speech quality metrics and subjective human judgments of coherence and similarity. Use when the user wants to benchmark on test-dialog-zh, test-dialog-en, or asks about evaluating this task. Reports cpWER.

researchpython
0
3
Zoombench EvalA

Evaluates fine-grained multimodal perception, visual grounding, and reasoning capabilities of vision-language models. It measures performance across general perception, specific perception (color, counting), and out-of-distribution generalization tasks using a suite of established and custom benchmarks. Use when the user wants to benchmark on ZoomBench, HR-Bench, VStar, CV-Bench, MME-RealWorld, ColorBench, CountQA, MMStar, BabyVision, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Zs Cir EvalA

Evaluates a model's ability to retrieve target images based on a reference image and a natural language modification text. It probes fine-grained visual-semantic alignment, compositional reasoning, and ranking precision under varying levels of distractors and semantic transformations. Use when the user wants to benchmark on CIRR, CIRCO, FashionIQ, GeneCIS, or asks about evaluating this task. Reports Recall@K.

researchpythongo
0
3
Zs Multimodal Ie EvalA

Evaluates zero-shot multimodal named entity typing and relation extraction. It probes a model's ability to align text and image modalities for fine-grained semantic recognition of unseen entity types and relations without additional training. Use when the user wants to benchmark on WikiDiverse, Zheng et al. MRE dataset, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Zs Xlt News Rec EvalA

Evaluates zero-shot cross-lingual news recommendation by measuring how effectively a model recommends articles in a target language to users who only consume news in a source language. It probes the model's ability to leverage multilingual sentence embeddings and click behavior fusion without task-specific fine-tuning on the target language. Use when the user wants to benchmark on MIND (small) / xMIND (small), or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
Gsm8k EvalA

Evaluate an LLM on GSM8K — 1K grade-school math word problems requiring 2-8 step arithmetic reasoning. Use when the user wants to measure math reasoning, mentions GSM8K, or asks "how good is my model at multi-step word problems?". Reports exact-match accuracy on the final numeric answer (parsed from "#### N" suffix).

researchpythongo
0
3
Mmmu EvalA

Evaluate a multimodal model (LMM) on MMMU — 11.5K college-level questions across 6 disciplines, 30 subjects, 30 image types (charts, MRI, music sheets, chemical structures...). Use when the user wants to benchmark a vision-language model's expert-level reasoning, mentions MMMU / MMMU-Pro, or asks "is my LMM at expert human level?". Reports micro-averaged accuracy.

businesspythongo
0
3
Ndcg At KA

Compute normalized Discounted Cumulative Gain at cutoff k (nDCG@k) — the standard ranking metric for retrieval / recommendation / search evaluation when relevance is graded. Use when the user has a list of (query, ranked_doc_ids, relevance_judgements) and wants to score the ranking quality, or mentions "nDCG / NDCG / DCG / ranking metric / IR metric / BEIR-style eval". Returns a number in [0, 1]; higher = better ranking.

researchpythongo
0
3
Pass At KA

Compute pass@k — the standard "any of N samples is correct" metric for code-generation evaluation (HumanEval / MBPP / LiveCodeBench / APPS / BigCodeBench / CodeContests). Use when the user has N samples per problem and wants the unbiased estimator of "probability at least one of the top-k is correct". Returns mean pass@k across the dataset, in [0, 1].

researchpython
0
3
Swe Bench EvalA

Evaluate a code-editing agent / LLM on SWE-Bench (real GitHub issues → patches that pass the PR's tests). Use when the user wants to benchmark a coding agent on realistic software-engineering tasks, mentions "SWE-Bench / SWE-Bench Lite / SWE-Bench Verified", or asks "can my model fix real GitHub issues?". Reports resolve_rate (% issues whose generated patch passes the original PR's hidden tests).

ai-agentspythongo
0
3
Future House Aviary Aviary Agent GymA

Aviary is FutureHouse's open-source gymnasium for defining and benchmarking LLM agents on scientific tasks (math, multi-hop QA, biological sequences, scientific literature search, Jupyter notebooks). Use when the user wants to evaluate an LLM agent on standardized scientific environments, build custom RL-style environments for agent training, or reproduce results from the Aviary paper.

ai-agentspythongo
0
3
Future House Edison Client Crow Literature QaA

Fast scientific literature Q&A with citations via FutureHouse's Crow agent (production PaperQA2). Use when the user wants a single, well-cited answer drawn from the published scientific literature — biology, chemistry, medicine, ML, etc. Handles one focused question per call. For multi-paper thematic synthesis use Falcon; for "has anyone done X" precedent queries use Owl.

researchpythongo
0
3
Future House Edison Client Falcon Deep LiteratureA

Deep, high-reasoning literature synthesis via FutureHouse's Falcon agent (LITERATURE_HIGH job). Use when the user wants a thematic review, gap analysis, or systematic synthesis across many papers — not a single fact lookup. Costs more credits and takes minutes longer than Crow but produces SOTA-quality scholarly output.

researchpythongo
0
3
Future House Edison Client Finch Data AnalysisA

Hosted biological-data-analysis agent (Finch) on the FutureHouse Platform. Hands a dataset + question to Finch, which builds a Jupyter notebook that explores, analyzes, and interprets the data. Use when the user has a biological dataset (omics, imaging, clinical) and a research question, and wants a multi-step analysis with code + results, not just a literature answer.

datapythongo
0
3
Future House Edison Client Owl Precedent SearchA

"Has anyone done this before?" — precedent search across the scientific literature via FutureHouse's Owl agent (formerly HasAnyone). Use when the user wants to know whether a specific experiment, technique, measurement, drug-target combination, or method has ever been published. Returns a yes/no-grounded answer with the closest matching prior work.

researchpythongo
0
3
Future House Edison Client Phoenix ChemistryA

Cheminformatics-grounded chemistry agent (Phoenix, the successor to ChemCrow) via the FutureHouse Platform. Use for retrosynthesis, reaction planning, molecular property prediction, SMILES manipulation, and proposing new molecules with chemistry tools backing the reasoning. Trigger on chemistry / drug-design / synthesis / molecule questions.

businesspythongo
0
3
Future House Ether0 Ether0 Chemistry RewardsA

Open-weights chemistry reasoning model + verifiable reward functions from FutureHouse's ether0 (arXiv 2506.17238). Use to score model-generated chemistry outputs (SMILES validity, molecular completion, synthesis reasoning) against ground truth, or to run the open-weights ether0 model itself for chemistry reasoning. Also useful for visualizing molecules and reactions from SMILES.

ai-agentspythongo
0
3
Future House Paper Qa Paperqa Local RagA

Run PaperQA2 locally on a folder of scientific PDFs to get high-accuracy, fully-cited answers. Self-hosted, open-source RAG (Apache-2.0) — needs only an LLM key (OpenAI/Anthropic/local), no FutureHouse credits. Use when the user has a local corpus of papers and wants grounded answers, or wants to avoid the hosted Crow/Falcon for privacy / cost reasons.

ai-agentspythonbash
0
3
Future House Paper Qa Wikicrow Article GeneratorA

Generate Wikipedia-style scientific articles section-by-section by orchestrating PaperQA2 over a topic-specific corpus. Reproduces the WikiCrow recipe used by FutureHouse to write the gene articles at wikicrow.ai. Use when the user wants a structured, fully-cited long-form article on a scientific topic (gene, protein, disease, drug, mechanism) rather than a single Q&A answer.

researchpythongo
0
3
Future House Robin Robin Disease DiscoveryA

Multi-agent automated scientific discovery for diseases — given a disease name, Robin generates and ranks experimental assays, proposes therapeutic candidates, and (optionally) analyzes wet-lab data. Open-source, Apache-2.0. Use when the user wants an end-to-end "I have a disease, give me hypotheses to test" workflow rather than a single literature lookup.

researchpythonbash
0
3
Imbad0202 Academic Research Skills Academic Paper ReviewerA

Multi-perspective academic paper review with dynamic reviewer personas. Simulates 5 independent reviewers (EIC + 3 peer reviewers + Devil's Advocate) with field-specific expertise. Supports full review, re-review (verification), quick assessment, methodology focus, Socratic guided, and calibration modes. Triggers on: review paper, peer review, manuscript review, referee report, review my paper, critique paper, simulate review, editorial review, calibrate reviewer, reviewer calibration, measur...

researchrustgo
0
3
Imbad0202 Academic Research Skills Academic PaperA

12-agent academic paper writing pipeline. 10 modes (full/plan/outline/revision/revision-coach/abstract/lit-review/format-convert/citation-check/disclosure). 6 paper types, 5 citation formats, bilingual abstracts, LaTeX/DOCX-via-Pandoc/PDF output. Style Calibration + Writing Quality Check + Anti-Patterns with IRON RULE markers. Triggers: write paper, academic paper, guide my paper, parse reviews, AI disclosure, 寫論文, 學術論文, 引導我寫論文, 審查意見.

researchpythongo
0
3
Imbad0202 Academic Research Skills Academic PipelineA

Orchestrator for the full academic research pipeline: research -> write -> integrity check -> review -> revise -> re-review -> re-revise -> final integrity check -> finalize. Coordinates deep-research, academic-paper, and academic-paper-reviewer into a seamless 10-stage workflow with mandatory integrity verification, two-stage peer review, and reproducible quality gates. Triggers on: academic pipeline, research to paper, full paper workflow, paper pipeline, end-to-end paper, research-to-publi...

researchgodocumentation
0
3
Imbad0202 Academic Research Skills Deep ResearchA

Universal deep research agent team. 13-agent pipeline for rigorous academic research on any topic. 7 modes: full research, quick brief, paper review, lit-review, fact-check, Socratic guided research dialogue, and systematic review with optional meta-analysis. Covers research question formulation, Socratic mentoring, methodology design, systematic literature search, source verification, cross-source synthesis, risk of bias assessment, meta-analysis, APA 7.0 report compilation, editorial review...

researchgoexpress
0
3
Antfu Skills Skills TurborepoA

Turborepo monorepo build system guidance. Triggers on: turbo.json, task pipelines, dependsOn, caching, remote cache, the "turbo" CLI, --filter, --affected, CI optimization, environment variables, internal packages, monorepo structure/best practices, and boundaries. Use when user: configures tasks/workflows/pipelines, creates packages, sets up monorepo, shares code between apps, runs changed/affected packages, debugs cache, or has apps/packages directories.

developmentjavascripttypescript
0
3