All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,215 views
Msvbench EvalA

Evaluates multi-shot video generation models on narrative coherence, cross-shot consistency, visual fidelity, and motion quality. It probes whether models can maintain character and scene identity across sequential shots and adhere to physical laws, rather than merely generating isolated visual interpolations. Use when the user wants to benchmark on MSVBench, or asks about evaluating this task. Reports Spearman’s ρ.

researchpythongo
0
3
Mt Bench EvalA

Evaluates the conversational quality and instruction-following capability of aligned language models across multiple knowledge domains. It also measures whether alignment fine-tuning causes regression in base reasoning, truthfulness, and commonsense capabilities. Use when the user wants to benchmark on MT-Bench, Open LLM Leaderboard Benchmarks, or asks about evaluating this task. Reports MT-Bench Average Score.

researchpythongo
0
3
Mt Data Filtering EvalA

This evaluation protocol assesses how effectively Quality Estimation (QE) metrics can filter low-quality or noisy sentence pairs from large parallel corpora. It measures whether retaining only the top 50% of high-scoring pairs improves downstream Neural Machine Translation (NMT) performance compared to using the full corpus or alternative filtering baselines like BICLEANER. Use when the user wants to benchmark on WMT & IWSLT Evaluation Campaigns, or asks about evaluating this task. Reports CO...

researchpythonperformance
0
3
Mt Geneval EvalA

Evaluates machine translation models' ability to correctly translate gender-specific words and maintain gender agreement across multiple languages. It also measures representational bias by comparing translation quality between male and female counterfactual sentence pairs. Use when the user wants to benchmark on MT-GenEval, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Mt Incremental EvalA

This benchmark evaluates how well automatic machine translation metrics track quality improvements in commercial systems over time. It probes whether metrics consistently rank newer systems higher than older ones, and how their reliability changes as system quality improves or when synthetic references are used. Use when the user wants to benchmark on Commercial MT Systems Corpus, or asks about evaluating this task. Reports Accuracy.

researchpythongit
0
3
Mt Orthographic Robustness EvalA

Evaluates the robustness of neural machine translation systems to orthographic and interpunctual noise by measuring translation quality and output consistency on perturbed inputs. Use when the user wants to benchmark on Baltic MT test sets (ET-EN, LV-EN, LT-EN), or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Mt Proxy CorrelationA

Evaluates whether machine translation quality serves as a scalable proxy for multilingual model performance on downstream tasks. It measures the alignment between MT metric scores and actual benchmark success across languages and model sizes. Use when the user has predictions and gold and needs to compute Pearson r.

researchpythongo
0
3
Mt Quality Estimation EvalA

Evaluates the ability of LLMs to predict human-assigned Direct Assessment (DA) scores for machine translation outputs. It probes how well different prompting strategies (zero-shot, CoT, few-shot) and input components (source, reference, error words) correlate with human judgments across various language pairs. Use when the user wants to benchmark on WMT QE / DA dataset (EN-DE, EN-MR, EN-ZH, ET-EN, NE-EN, RO-EN, RU-EN, SI-EN), or asks about evaluating this task. Reports Spearman $ ho$.

researchpythongo
0
3
Mt Raig EvalA

This benchmark evaluates retrieval-augmented insight generation over multiple tables. It requires models to retrieve relevant tables from a database and synthesize multi-hop insights across them, assessing both faithfulness and completeness of the generated reasoning. Use when the user wants to benchmark on MT-RAIG Bench, or asks about evaluating this task. Reports MT-RAIG Eval.

researchpythongo
0
3
Mt Reasoning Scaling EvalA

This evaluation probes the test-time scaling properties of reasoning models across diverse machine translation tasks. It measures how varying reasoning budgets and iterative self-correction workflows impact translation quality across literary, biomedical, cultural, and commonsense domains. Use when the user wants to benchmark on WMT24-Literary, MetaphorTrans, LitEval-Corpus, WMT24-Biomedical, WMT23-Biomedical, CAMT, Commonsense-MT, RTT, RAGTrans, or asks about evaluating this task. Reports CO...

researchpythonexpress
0
3
Mtass Sdr EvalA

Evaluates a model's ability to simultaneously separate speech, music, and background noise from monaural audio mixtures. It measures separation fidelity using signal-to-distortion ratio and its improvement over baselines across all three source tracks. Use when the user wants to benchmark on Unspecified (speech, music, noise tracks), or asks about evaluating this task. Reports SDRi (dB).

researchpythonperformance
0
3
Mtat Nsynth Fma EvalA

Evaluates the quality of self-supervised audio representations on three downstream tasks: music auto-tagging, instrument family classification, and music genre classification. It probes how well contrastive learning objectives capture semantic and structural musical features. Use when the user wants to benchmark on MTAT, NSynth, FMA, or asks about evaluating this task. Reports ROC-AUC.

researchpythongo
0
3
Mtbbench EvalA

Evaluates AI agents' ability to perform longitudinal, multimodal clinical decision-making in oncology. Agents must integrate evolving patient data across pathology, genomics, hematology, and imaging over multiple turns to answer diagnostic and prognostic questions, simulating molecular tumor board workflows. Use when the user wants to benchmark on MTBBench-Multimodal (HANCOCK subset), MTBBench-Longitudinal (MSK-CHORD subset), or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Mtbench EvalA

Evaluates the model's ability to generate harmless and helpful responses in multi-turn conversations, measuring the trade-off between safety alignment and utility. Use when the user wants to benchmark on MTBench, or asks about evaluating this task. Reports MTBench harmlessness score.

researchpythongo
0
3
Mtbi Speech EvalA

Evaluates the automatic speech recognition accuracy and zero-shot generalization capabilities of Speech Large Language Models. It probes robustness across mathematical reasoning, speaker role inference, and prompt adaptation tasks. Use when the user wants to benchmark on LibriSpeech, GSM8K, Generalization Test Set, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Mtc FragmentsA

Compute mtc/fragments via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mtc/fragments.

developmentpython
0
3
Mtc Locomotion EvalA

Probes a humanoid robot's ability to navigate procedurally generated 3D cluttered environments while adapting full-body kinematics to geometric constraints. It quantifies how much a policy deviates from nominal flat-ground walking and measures collision safety against complex scene geometry. Use when the user wants to benchmark on MTC Dataset, or asks about evaluating this task. Reports Motion Adaptation Score.

researchpython
0
3
Mtcityscapes 3d EvalA

Evaluates joint 2D-3D multi-task scene understanding on urban street imagery. It probes a model's ability to concurrently perform monocular 3D vehicle detection, 19-class semantic segmentation, and monocular depth estimation. Use when the user wants to benchmark on MTCityscapes-3D, or asks about evaluating this task. Reports mDS.

researchpythongo
0
3
Mteb Airbench EvalA

Evaluates text embedding models on retrieval, reranking, clustering, classification, semantic textual similarity, and summarization tasks. It also assesses out-of-domain generalization on domain-specific question answering and retrieval benchmarks. Use when the user wants to benchmark on MTEB, AIR-Bench, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythonperformance
0
3
Mteb Beir Miracl EvalA

Evaluates text embedding models across diverse NLP tasks including classification, clustering, semantic textual similarity, reranking, and information retrieval. It specifically probes multilingual retrieval capabilities and measures how effectively synthetic data generation improves embedding quality without relying on labeled supervision. Use when the user wants to benchmark on MTEB (English subset), BEIR (Retrieval), MIRACL, or asks about evaluating this task. Reports MTEB official metrics.

ai-agentspythongo
0
3
Mteb Clustering EvalA

Evaluates the quality of text embeddings for document clustering by measuring how well the embeddings group semantically related sentences together. It probes the model's ability to capture semantic similarity without task-specific fine-tuning or with lightweight adaptation. Use when the user wants to benchmark on MTEB v1.38.30 English Clustering, or asks about evaluating this task. Reports clustering accuracy.

ai-agentspython
0
3
Mteb Echo EvalA

Evaluates the quality of text embeddings extracted from autoregressive language models in zero-shot and fine-tuned settings across a broad suite of downstream NLP tasks including classification, clustering, retrieval, and semantic textual similarity. Use when the user wants to benchmark on MTEB, or asks about evaluating this task. Reports Average MTEB Score.

ai-agentspythongo
0
3
Mteb Eng V2 EvalA

Evaluates the semantic discrimination and generalization capabilities of text embedding models across diverse NLP tasks including retrieval, classification, clustering, and semantic similarity. It uses a zero-shot English-only benchmark to measure performance without task-specific fine-tuning. Use when the user wants to benchmark on MTEB(eng, v2), or asks about evaluating this task. Reports average score across all tasks.

ai-agentspythongo
0
3
Mteb English EvalA

Evaluates the ability of decoder-only LLMs to generate universal text embeddings across diverse natural language processing tasks. It probes retrieval, reranking, clustering, classification, pair classification, semantic textual similarity, and summarization capabilities using standardized benchmark datasets. Use when the user wants to benchmark on MTEB (English subset), or asks about evaluating this task. Reports Average score.

ai-agentspythongo
0
3
Mteb EvalA

This evaluation probes the ability of decoder-only LLMs to generate high-quality, fixed-dimensional text embeddings without fine-tuning. It measures semantic similarity, information retrieval, classification, clustering, and long-context comprehension across diverse tasks and sequence lengths. Use when the user wants to benchmark on MTEB, LoCoV1, or asks about evaluating this task. Reports MTEB Average Score.

ai-agentspythongo
0
3
Mteb Loco Jina EvalA

Evaluates text embedding models on short and long-context retrieval, clustering, and semantic similarity tasks to measure representation quality across varying sequence lengths. Use when the user wants to benchmark on MTEB, Jina Long Context Benchmark, LoCo Benchmark, or asks about evaluating this task. Reports NDCG@10.

ai-agentspythongo
0
3
Mteb Longembed EvalA

Evaluates text embedding models on English, multilingual, and long-context retrieval tasks to measure semantic similarity, classification, clustering, reranking, and retrieval performance. Use when the user wants to benchmark on MTEB(eng, v2), MTEB(Multilingual, v2), LongEmbed, or asks about evaluating this task. Reports mean over tasks.

ai-agentspythongo
0
3
Mteb Mmteb Retrieval EvalA

Evaluates text embedding models across retrieval, semantic similarity, clustering, classification, and reranking tasks. It probes the ability of compact, distillation-trained models to generalize across multilingual corpora, long documents, and enterprise-scale retrieval benchmarks. Use when the user wants to benchmark on MTEB (English v2), MMTEB (Multilingual v2), RTEB (Multilingual), BEIR, LongEmbed, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythongo
0
3
Mteb Negation EvalA

Evaluates the semantic similarity and retrieval capabilities of sentence embedding models across diverse downstream tasks. It also probes the models' sensitivity to grammatical negation and their ability to distinguish syntactically similar negative examples from entailments. Use when the user wants to benchmark on MTEB benchmark, Negation dataset, or asks about evaluating this task. Reports nDCG@10.

ai-agentspythongo
0
3
Mteb Retrieval EvalA

Evaluates text embedding models on information retrieval tasks using the MTEB benchmark and a custom e-commerce Q&A dataset. It measures ranking quality via nDCG and mAP, and assesses similarity distribution calibration via AUPRC on a held-out set with single-relevant-passage queries. Use when the user wants to benchmark on MTEB Retrieval, E-commerce Q&A, or asks about evaluating this task. Reports nDCG@10.

researchpythonperformance
0
3
Mteb Subset EvalA

Evaluates text embedding models across diverse semantic tasks including retrieval, reranking, clustering, pair classification, classification, and semantic textual similarity to measure the quality of dense vector representations. Use when the user wants to benchmark on MTEB (15-task subset), or asks about evaluating this task. Reports MTEB average score.

ai-agentspythongo
0
3
Mthl Network Traffic EvalA

Evaluates machine learning models on hierarchical network traffic classification tasks, including top-level protocol identification and malware detection, as well as mid-level application and malware type classification. Use when the user wants to benchmark on VPN-nonVPN + $(\mathsf{Net})^2$ + CICIDS2017, or asks about evaluating this task. Reports Macro-average F1 score.

datapythongit
0
3
Mti Bench EvalA

Evaluates whether large language models can process multiple distinct instructions simultaneously within a single inference call, compared to sequential or batched approaches. It probes reasoning consistency, format adherence, and inference efficiency across a diverse set of 28 NLP tasks. Use when the user wants to benchmark on MTI Bench, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Mti Temperament EvalA

Measures four independent behavioral axes—Reactivity, Compliance, Sociality, and Resilience—in AI agents to distinguish intrinsic dispositional traits from raw capability. It evaluates how models respond to structured behavioral protocols under baseline and stress conditions, and how alignment techniques like RLHF alter these dispositions. Use when the user wants to benchmark on MTI Behavioral Battery, or asks about evaluating this task. Reports MTI Temperament Axes (Reactivity, Compliance, S...

researchpythonshell
0
3
Mtop EvalA

Evaluates multilingual task-oriented semantic parsing models on extracting hierarchical intent-slot representations from natural language utterances across multiple languages and transfer settings. It tests the model's ability to generalize across languages using in-language, multilingual, and zero-shot training protocols. Use when the user wants to benchmark on MTOP, Multilingual ATIS, Multilingual TOP, or asks about evaluating this task. Reports exact match.

researchpythongo
0
3
Mtqe Generation Based EvalA

Evaluates machine translation quality estimation (MTQE) methods by measuring how well their segment-level scores correlate with human judgments across multiple language pairs. It specifically tests a generation-based paradigm where LLMs create reference translations instead of directly scoring outputs. Use when the user wants to benchmark on WMT22 Test Sets (8 language pairs), or asks about evaluating this task. Reports Spearman rank correlation (ρ).

researchpython
0
3
Mtr Duplexbench EvalA

This benchmark evaluates Full-Duplex Speech Language Models (FD-SLMs) on their ability to sustain performance across multi-round conversations. It probes dialogue quality, conversational dynamics (turn-taking, interruptions, pauses, background speech), instruction following, and safety, specifically measuring how these capabilities degrade or hold up as interaction rounds increase. Use when the user wants to benchmark on MTR-DuplexBench, Llama Question, AdvBench, or asks about evaluating this...

researchpythongo
0
3
Mtrag Un EvalA

Evaluates multi-turn RAG systems on handling unanswerable, underspecified, and non-standalone questions. It probes both retrieval ranking quality and generation quality, including the model's ability to correctly refuse or request clarification when context is insufficient. Use when the user wants to benchmark on MTRAG-UN, or asks about evaluating this task. Reports RB_llm.

ai-agentspythongo
0
3
Mtvqa EvalA

This benchmark evaluates the multilingual visual-textual alignment and comprehension capabilities of multimodal large language models (MLLMs). It specifically probes whether models can accurately perceive, extract, and reason about text embedded within images across nine different languages without relying on translation. Use when the user wants to benchmark on MTVQA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Mtzig Cross Entropy LossA

Compute mtzig/cross_entropy_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mtzig/cross_entropy_loss.

developmentpython
0
3
Muben Uq EvalA

This benchmark evaluates the uncertainty quantification (UQ) capabilities of molecular representation models across diverse backbone architectures and input modalities. It probes how accurately models predict molecular properties (binary classification and continuous regression) while simultaneously estimating their own predictive uncertainty under both in-distribution and out-of-distribution conditions. The evaluation specifically tests robustness to molecular scaffold shifts, which better s...

researchpythonperformance
0
3
Much EvalA

Evaluates the ability of logit-based uncertainty quantification (UQ) methods to predict claim-level hallucination (factuality) in multilingual LLM outputs. It measures how well token-level confidence scores, when aggregated, correlate with ground-truth factuality labels across different languages and model configurations. Use when the user wants to benchmark on MUCH, or asks about evaluating this task. Reports ROC-AUC.

researchpythongo
0
3
Muchin EvalA

Evaluates language models' ability to generate structured Chinese lyrics from music descriptions and to understand music audio by generating descriptive tags. It probes alignment with public (amateur) vs. professional musical perception and semantic similarity in Chinese. Use when the user wants to benchmark on MuChin, or asks about evaluating this task. Reports Overall Score.

researchpythongo
0
3
Muchomusic EvalA

Evaluates multimodal audio-language models' ability to understand music through factual knowledge and reasoning tasks. It probes whether models can ground their answers in audio content rather than relying on language priors or hallucinating musical elements. Use when the user wants to benchmark on MuChoMusic, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
MucrevalA

Evaluates vision-language models' ability to infer causal relationships across text and image modalities using siamese image-text pairs. It probes cross-modal generalization and visual cue identification in causal reasoning tasks. Use when the user wants to benchmark on MuCR, or asks about evaluating this task. Reports C2E score.

researchpythongo
0
3
Mucue EvalA

Evaluates a model's ability to understand music across a spectrum of tasks, ranging from low-level acoustic perception (e.g., pitch, chord, rhythm) to high-level cognitive reasoning (e.g., genre, mood, structure, lyrical comprehension). It probes whether foundation models can process long-context audio and lyrics jointly to answer standardized multiple-choice questions. Use when the user wants to benchmark on MuCUE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mudabench EvalA

MuDABench probes multi-document analytical QA capabilities, requiring models to synthesize quantitative insights and perform inter-document reasoning across large collections of heterogeneous financial documents. It specifically tests cross-document filtering, conditional information extraction, and multi-step numerical computation beyond standard single-document retrieval. Use when the user wants to benchmark on MuDABench, or asks about evaluating this task. Reports final-answer accuracy.

researchpythongo
0
3
Mudaif Vl EvalA

Evaluates a decoder-only vision-language model's ability to perform visual question answering, image captioning, and multimodal reasoning. It measures cross-modal alignment, computational efficiency, and robustness to input variations like resolution and noise. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, SEED, MM-Vet, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Mudi EvalA

This benchmark evaluates a model's ability to predict pharmacodynamic drug-drug interactions (Synergism, Antagonism, or New Effect) using multimodal inputs including text, chemical formulas, molecular graphs, and images. It specifically probes cross-modal reasoning and generalization to unseen drug pairs under both direction-aware and direction-agnostic matching settings. Use when the user wants to benchmark on MUDI, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Muennighoff Code Eval OctopackA

Compute Muennighoff/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Muennighoff/code_eval_octopack.

developmentpython
0
3