
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the accuracy of a joint multiphysics-decision tree learning framework in estimating subsurface transport parameters and simulating state variables (pressure head, temperature, concentration) under stochastic boundary conditions. It benchmarks the reduced-order surrogate models against a full numerical multiphysics inversion baseline. Use when the user wants to benchmark on Stochastic managed aquifer recharge dataset, or asks about evaluating this task. Reports Nash-Sutcliffe efficie...
Evaluates large language models' ability to generate functionally correct code in low-resource programming languages (R and Racket). It probes how well in-context learning strategies and fine-tuning adapt pre-trained models to languages with limited training data and documentation. Use when the user wants to benchmark on MultiPL-E, or asks about evaluating this task. Reports pass@1.
This evaluation probes the true multimodal reasoning capability of vision-language models on multiple-choice question answering tasks. It specifically measures whether models rely on actual question understanding or exploit visual relevance imbalances between correct answers and distractors. Performance is assessed under both standard (vision, question, options) and question-omitted (vision, options) settings to detect easy-option bias. Use when the user wants to benchmark on NExT-QA, MMStar,...
Evaluates the multilingual language fidelity and question-answering accuracy of open LLMs across 137 typologically diverse languages. It probes whether models respond in the prompt's language and whether their answers are factually correct, highlighting the impact of tokenization strategies and model scaling on multilingual performance. Use when the user wants to benchmark on MultiQ, or asks about evaluating this task. Reports QA accuracy (%).
Evaluates a model's ability to perform real-time, multimodal sequence labeling to detect and classify questions in emergency call speech. It probes robustness to noisy ASR transcriptions and temporal alignment under streaming conditions. Use when the user wants to benchmark on question and symptoms tracking datasets, or asks about evaluating this task. Reports TIMESTEP F1.
Evaluates the ability of image generation models to simultaneously align and incorporate multiple visual reference conditions (e.g., bounding boxes, depth maps, masks, sketches) alongside text instructions. It probes complex multi-source creative synthesis, testing both global image quality and fine-grained reference fidelity across different input formats and processing orders. Use when the user wants to benchmark on MULTIREF-BENCH, or asks about evaluating this task. Reports Overall Assessm...
Evaluates machine learning model robustness against multiple diverse adversarial attacks (e.g., ℓₚ-norm, color shifts, spatial transformations) across varying strengths. It quantifies how well defenses maintain performance under worst-case and average-case multiattack scenarios, addressing bias from varying attack difficulties. Use when the user wants to benchmark on MultiRobustBench, or asks about evaluating this task. Reports competitiveness ratio (CR).
Compute the MultiScaleStructuralSimilarityIndexMeasure metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultiScaleStructuralSimilarityIndexMeasure, or asks how to score with MultiScaleStructuralSimilarityIndexMeasure.
Evaluates multi-person spatio-temporal action detection in sports videos, probing the model's ability to localize fine-grained actions across multiple concurrent persons, handle occlusion, and model long-range temporal context. Use when the user wants to benchmark on MultiSports, or asks about evaluating this task. Reports frame-mAP@0.5.
Evaluates a fine-tuned LLM's ability to mitigate toxicity while preserving general knowledge and utility across multiple benchmarks. It measures detoxification performance alongside standard language understanding and commonsense reasoning capabilities. Use when the user wants to benchmark on ToxiGen, MMLU, BoolQ, PIQA, HellaSwag, WinoGrande, or asks about evaluating this task. Reports MMLU (utility).
Compute the MultitaskWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultitaskWrapper, or asks how to score with MultitaskWrapper.
Evaluates multimodal models' ability to understand and interact with complex webpage UIs, perform text-rich visual grounding, and generalize to OCR and document understanding tasks. Use when the user wants to benchmark on VisualWebBench, Mind2Web, DocVQA, ChartQA, or asks about evaluating this task. Reports element accuracy.
Evaluates the ability of spatiotemporal attention models to accurately predict future values in multivariate time series across environmental, building HVAC, and clinical domains. It also probes the model's capacity to produce interpretable attention weights that align with known physical or physiological relationships. Use when the user wants to benchmark on Beijing PM2.5 Data Set, Building HVAC Dataset, MIMIC-III EHR Dataset, or asks about evaluating this task. Reports RMSE.
Evaluates event-centric video retrieval across six languages, requiring models to match natural language queries about specific world events to relevant long-form videos using multimodal signals (vision, audio, OCR, metadata). It probes a model's ability to integrate cross-lingual, cross-modal information for complex event understanding rather than simple visual matching. Use when the user wants to benchmark on MultiVENT 2.0, or asks about evaluating this task. Reports Retrieval Performance.
Evaluates the multi-turn conversational reasoning and sustained dialogue capabilities of Vision-Language Models (VLMs) across diverse domains like mathematics, coding, and creative tasks. It probes how well models leverage dialogue history (in-context learning) and maintain consistency over extended interactions. Use when the user wants to benchmark on MultiVerse, or asks about evaluating this task. Reports checklist-based evaluation.
This benchmark evaluates a model's ability to perform product-level composed image retrieval (CIR) in fashion e-commerce, specifically handling multi-view product images and short modification queries. It probes the model's capacity to align visual perception with textual reasoning across multiple views while filtering out irrelevant gallery items. Use when the user wants to benchmark on DeepFashion, Fashion200K, FashionGen-val, or asks about evaluating this task. Reports Recall@5.
Evaluates voice assistants' ability to jointly ground visual and paralinguistic speech cues (e.g., pitch, emotion, volume, background sounds) in context-aware responses. It specifically tests robustness against confounding samples that flip speech properties to prevent overreliance on unimodal priors. Use when the user wants to benchmark on MultiVox, or asks about evaluating this task. Reports visual grounding and non-verbal speech signals.
Evaluates a model's ability to track and predict the complete set of user intent slots (dialogue state) across multiple domains in a multi-turn conversation. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports slot accuracy.
Evaluates a model's ability to perform Dialogue State Tracking (DST) by predicting the correct values and statuses for all requested slots across multi-domain conversations. It specifically probes robustness to long-range contextual noise and class imbalance in slot status prediction. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint-goal-accuracy (JGA).
Evaluates abstractive summarization models on their ability to preserve critical semantic slots, entities, and domain consistency across multi-domain dialog conversations. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports ROUGE.
Evaluates the quality, diversity, and goal adherence of task-oriented dialogue generation models. It measures how well a model generates natural, diverse responses while correctly incorporating specified dialogue goals and slot values. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports BLEU-4.
Evaluates a model's ability to track and predict dialogue states across multiple domains in a conversation. It measures how accurately the system maintains slot-value pairs as the user's goals evolve and switches between domains like restaurant, hotel, and taxi. Use when the user wants to benchmark on MultiWOZ, or asks about evaluating this task. Reports Joint Goal Accuracy (JGA).
Evaluates end-to-end task-oriented dialogue systems on their ability to track user goals, fulfill multi-domain requests, and generate contextually appropriate responses. It specifically probes how well models maintain conversation state and achieve user objectives without relying on full historical dialogue context. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports inform rate, success rate.
Evaluates a model's ability to track dialogue state (domain, slot, value triplets) across conversation turns, specifically probing its robustness to user mind-changes or 'turnback' utterances that modify previously stated intentions. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint goal accuracy.
Evaluates a model's ability to perform multi-objective molecular lead optimization by modifying a starting molecule to simultaneously improve multiple conflicting pharmacological properties while retaining structural similarity. Use when the user wants to benchmark on MuMO-Instruct, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates multimodal large language models' ability to generate verifiable, fact-level citations grounded in video and audio inputs. It probes whether models can correctly decompose reasoning into atomic claims and align them with precise temporal and modality-specific evidence without hallucinating references. Use when the user wants to benchmark on Video-MMMU, WorldSense, or asks about evaluating this task. Reports MURGAT-S.
Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks. Use when the user wants to benchmark on Multilingual MMLU, TranslatedDolly, Taxi1500, or asks about evaluating this task. Reports accuracy.
Compute murinj/hter via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of murinj/hter.
Evaluates large language models' ability to perform multi-step, commonsense-rich reasoning over long natural language narratives. It probes whether models can follow complex, implicit logical chains (e.g., murder motives, object spatial reasoning, team skill matching) without relying on simple keyword heuristics or rule-based shortcuts. Use when the user wants to benchmark on MuSR, or asks about evaluating this task. Reports accuracy.
Evaluates the computational efficiency and alignment quality of a multiple sequence alignment algorithm across genomic and protein datasets. It probes the trade-off between runtime scalability and evolutionary accuracy metrics like distance distortion and gap percentage. Use when the user wants to benchmark on Greengenes 12.10, Greengenes 13.5, PDB, PFam-10k, PFam-100k, PFam-1M, or asks about evaluating this task. Reports runtime, distance distortion.
Evaluates multilingual automatic speech recognition (ASR) systems on spontaneous scientific conversations, focusing on their ability to handle code-switching, transcribe domain-specific technical terms, and maintain accuracy across varying audio recording devices and segmentation methods. Use when the user wants to benchmark on MUSCAT, or asks about evaluating this task. Reports WER.
Multimodal scientific claim verification, requiring models to read complex figures and captions to determine if a scientific claim is supported, neutral, or contradicted. It also probes evidence localization, basic visual understanding, cross-modal aggregation, and epistemic sensitivity. Use when the user wants to benchmark on MuSciClaims, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to isolate individual musical stems (vocals, drums, bass, other) from mixed audio recordings, testing long-range context modeling and cross-domain attention capabilities in source separation. Use when the user wants to benchmark on MUSDB, or asks about evaluating this task. Reports SDR.
Evaluates the ability of waveform-to-waveform models to separate individual musical instruments (drums, bass, other, vocals) from a mixed audio track. It probes the model's capacity to isolate sources while minimizing contamination and artifacts, measured against ground-truth stems. Use when the user wants to benchmark on MusDB, or asks about evaluating this task. Reports SDR.
Evaluates the perceptual quality of multi-stem music source separation (vocals, drums, bass, other) generated by a discrete token modeling framework. It measures how well the model separates audio tracks compared to discriminative baselines, focusing on perceptual audio quality and vocal intelligibility/naturalness. Use when the user wants to benchmark on MUSDB18-HQ, or asks about evaluating this task. Reports ViSQOL.
Evaluates the quality of separated audio sources (vocals, drums, bass, other) from mixed music tracks using source-to-distortion ratio, probing the model's ability to perform supervised music source separation. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.
Probes a model's ability to isolate specific audio sources from a mixture using a provided query signal. It evaluates how well the model handles continuous latent-space conditioning and separates arbitrary or subclass instruments beyond standard training labels. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.
Evaluates an LLM-based planning framework's ability to decompose natural language queries into correct task selections, logical execution flows, and valid final multimodal outputs. It probes constraint-aware model orchestration and multi-modal task routing across heterogeneous AI services. Use when the user wants to benchmark on MuSE, or asks about evaluating this task. Reports Task Selection (TS).
Evaluates multimodal language models' ability to perform fine-grained, interactive reasoning over symbolic music scores and expressive performance audio. It probes capabilities in score–audio alignment, performance error detection, and expressive deviation analysis across text, audio, and image modalities. Use when the user wants to benchmark on MuseBench, or asks about evaluating this task. Reports Accuracy (%).
This benchmark evaluates the zero-shot instance segmentation capability of models trained on synthetic data when applied to real-world agricultural scenes. It probes the model's ability to generalize across domain gaps, handling varying lighting, mushroom densities, and developmental stages without fine-tuning on real annotations. Use when the user wants to benchmark on Real-world on-field dataset, M18K, or asks about evaluating this task. Reports F1 score.
Evaluates the quality of pre-trained audio embeddings for downstream music understanding tasks including tagging, genre classification, mood prediction, pitch/instrument detection, key classification, and emotion recognition. It tests whether frozen embeddings can be effectively probed with simple MLP classifiers to achieve competitive performance without fine-tuning the backbone model. Use when the user wants to benchmark on MSDS, MSD50, MSD100, MSD500, AMM, MuMu, MTT, NSynthP, NSynthI, GTZA...
This evaluation probes a model's ability to perform large-scale music audio tagging by predicting a fixed set of semantic labels (e.g., genre, mood, instrumentation) from raw 30-second audio clips. It measures how well architectures generalize across varying dataset sizes and label granularities. Use when the user wants to benchmark on MagnaTagATune (MTT), Million Song Dataset (MSD), Private Dataset, or asks about evaluating this task. Reports top-50 tag prediction.
Evaluates audio representation models on music autotagging tasks, measuring how well they predict categorical genre/instrument/mood tags and continuous musical features from audio input. It compares performance across generic tag datasets and expert-annotated continuous features to highlight limitations in current evaluation practices. Use when the user wants to benchmark on MagnaTagATune, MTG-Jamendo, MGPHot-tag, MGPHot-reg, or asks about evaluating this task. Reports MAP.
Evaluates a model's ability to generate fine-grained, temporally-aware music captions and retrieve corresponding audio segments using those captions. It probes the model's capacity for temporal reasoning, structural music understanding, and cross-modal alignment. Use when the user wants to benchmark on MusicCaps, Song Descriptor, or asks about evaluating this task. Reports BLEU-1/2/3, METEOR, ROUGE-L, BERTScore, Recall@K, Median Rank.
This evaluation probes a diffusion-based music generation model's ability to precisely follow time-varying control signals (melody, dynamics, rhythm) and global style tags (genre/mood). It measures how faithfully the generated audio adheres to these inputs while maintaining overall audio realism and diversity. Use when the user wants to benchmark on In-domain test set, MusicCaps, MusicCaps+ChatGPT, Created Controls dataset, or asks about evaluating this task. Reports Melody accuracy.
Evaluates a model's ability to detect plagiarized or remixed segments within audio tracks by computing segment-level musical similarity and attributing similarities to specific elements like melody, chords, and vocals. Use when the user wants to benchmark on Similar Music Pair, or asks about evaluating this task. Reports similarity score.
Evaluates zero-shot language-queried audio source separation on musical instrument classes. The benchmark tests the model's ability to isolate a target instrument from a mixed audio mixture using text labels. Use when the user wants to benchmark on MUSIC, or asks about evaluating this task. Reports SDRi.
Evaluates a model's ability to predict multiple audio tags (e.g., genre, mood, instruments) from short audio segments. It probes long-range temporal dependency modeling and robustness to class imbalance in user-generated music metadata. Use when the user wants to benchmark on MagnaTagATune (MTAT), Million Song Dataset (MSD), or asks about evaluating this task. Reports AUPR.
Evaluates a model's ability to generate high-fidelity, long-form music from complex text descriptions. It probes both audio quality/plausibility and the model's adherence to specific textual constraints such as genre, mood, tempo, and instrumentation. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.
Evaluates the capability of text-to-music generation models to produce high-fidelity, controllable audio that aligns with textual descriptions and matches human perceptual quality standards. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.