
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the ability of text-to-image generative models to produce visually coherent and structurally plausible music score images conditioned on textual descriptions of musical attributes like instrumentation, key, and composer. It benchmarks visual fidelity and distribution matching against ground-truth sheet music. Use when the user wants to benchmark on MusicScore-400, MusicScore-14k, MusicScore-200k, or asks about evaluating this task. Reports FID.
Evaluates multimodal models on their ability to understand, generate, and retrieve music based on semantically rich, context-aware natural language descriptions. It probes fine-grained musical semantics beyond technical attributes, including atmospheric, situational, and contextual cues. Use when the user wants to benchmark on MusicSem, or asks about evaluating this task. Reports BLEU.
Evaluates a model's ability to reason about music theory concepts and understand symbolic music representations, alongside general language knowledge and structured music generation capabilities. Use when the user wants to benchmark on MusicTheoryBench, MMLU, or asks about evaluating this task. Reports average accuracy.
Evaluates the quality of symbolic music generation by spiking neural networks across multiple datasets. It assesses both objective statistical properties (pitch, rhythm, harmony) and subjective cognitive/perceptual dimensions (fluency, emotion, impression, autobiographical association). Use when the user wants to benchmark on JSB Chorales, POP909, Lakh MIDI, EMOPIA, XMIDI, or asks about evaluating this task. Reports Personal preference.
This evaluation probes a model's ability to answer music-specific factual and contextual questions using retrieval-augmented generation. It measures accuracy on both in-domain artist metadata and out-of-domain music knowledge across multiple-choice formats. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.
Compute the mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute mutual_info_score, or asks how to score with mutual_info_score.
Compute the MutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MutualInfoScore, or asks how to score with MutualInfoScore.
Evaluates a model's ability to perform multi-task view synthesis by predicting multiple scene properties (RGB, surface normals, shading, edges, keypoints, semantic segmentation) from novel viewpoints, given a set of source-view annotations and camera poses. Use when the user wants to benchmark on Replica, SceneNet RGB-D, or asks about evaluating this task. Reports RGB.
Evaluates a diffusion adapter's ability to generate geometrically consistent multi-view images conditioned on text prompts or reference images with camera parameters. It measures visual fidelity, image-text alignment, and multi-view structural similarity against ground-truth 3D scans. Use when the user wants to benchmark on Objaverse, Google Scanned Objects (GSO), or asks about evaluating this task. Reports FID.
Evaluates multi-modal large language models' ability to understand video content, with a strong focus on temporal perception and static-to-dynamic task transformation across 20 diverse categories ranging from basic perception to complex reasoning. Use when the user wants to benchmark on MVBench, or asks about evaluating this task. Reports accuracy.
Evaluates cross-modal and text-only topical matching capabilities of vision-language models across 205 languages. It probes whether models can correctly associate images with semantically related texts (or vice versa) in a multilingual multiple-choice setting. Use when the user wants to benchmark on MVL-SIB, or asks about evaluating this task. Reports accuracy.
Evaluates text classification models on sentiment analysis and news categorization tasks. It probes the model's ability to aggregate diverse feature views (word-level and n-gram) to predict fine-grained sentiment categories and news topics. Use when the user wants to benchmark on Stanford Sentiment Treebank, AG News, or asks about evaluating this task. Reports accuracy.
This evaluation probes a model's ability to generate high-quality, multi-view consistent 3D textures on arbitrary meshes conditioned on text instructions. It measures visual fidelity, distributional similarity to ground truth, and cross-view consistency through both automated generative metrics and human preference studies. Use when the user wants to benchmark on Objaverse T2T benchmark, GSO T2T benchmark, or asks about evaluating this task. Reports FID.
Evaluates a vision-language model's few-shot classification capability on diverse biomedical images, testing cross-modal alignment, generalization to unseen disease categories, and robustness across multiple imaging modalities and anatomical regions. Use when the user wants to benchmark on CTKidney, DermaMNIST, Kvasir, RETINA, LC25000, CHMNIST, BTMRI, OCTMNIST, BUSI, COVID-QU-Ex, KneeXray, or asks about evaluating this task. Reports classification accuracy (%).
Evaluates an image classification and segmentation model's ability to detect and localize defects in industrial products without seeing anomalous examples during training. It probes the model's capacity to learn nominal feature distributions and identify deviations at both image and pixel levels. Use when the user wants to benchmark on MVTec AD, Magnetic Tile Defects (MTD), Mini Shanghai Tech Campus (mSTC), or asks about evaluating this task. Reports AUROC.
Unsupervised anomaly detection and pixel-level localization on industrial defect data. It probes the model's ability to distinguish normal from defective samples and precisely segment defect regions without using labeled anomalies during training. Use when the user wants to benchmark on MVTec, or asks about evaluating this task. Reports AUROC.
Evaluates an algorithm's ability to optimally partition repair targets among multiple crews and route them to minimize total weighted latency (average wait time) while balancing workload distribution across crews in post-disaster urban scenarios. Use when the user wants to benchmark on Random Environments, Champaign Case Study, or asks about evaluating this task. Reports wait.
Evaluates LLMs' ability to solve math word problems after socio-cultural localization of entities into low-resource languages. Probes whether models maintain reasoning accuracy when cultural context shifts, focusing solely on final answer correctness rather than step-by-step reasoning. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates mathematical reasoning and robustness on single-equation math word problems. It probes a model's ability to parse linguistic variations, ignore irrelevant information, and solve inverted or structurally complex problems. Use when the user wants to benchmark on MAWPS, SVAMP, PARAMAWPS, or asks about evaluating this task. Reports Value accuracy.
Evaluates the raw execution speed, memory footprint, and distributed scalability of the MXNet deep learning framework against Torch7, Caffe, and TensorFlow. It measures how efficiently the library handles standard convolutional neural network architectures and large-scale image classification tasks across single and multiple GPU nodes. Use when the user wants to benchmark on convnet-benchmarks, ILSVRC12, or asks about evaluating this task. Reports forward-backward performance.
Evaluates active few-shot learning frameworks for histopathology image classification under extremely tight annotation budgets (1, 5, and 10 labeled samples). It probes how well uncertainty-based diversity sampling and self-supervised contrastive pretraining can reduce sample redundancy and improve classification performance compared to standard few-shot and active learning baselines. Use when the user wants to benchmark on NCT-CRC-HE-100K, BreaKHis, or asks about evaluating this task. Report...
Evaluates clinical NLP models on extracting medical concepts and their relations from clinical notes. It probes the model's ability to handle nested/overlapped concepts and assesses cross-institutional generalization across different benchmark years. Use when the user wants to benchmark on n2c2 2018, n2c2 2022, n2c2 cross-institution (MIMIC-train/UW-test), or asks about evaluating this task. Reports strict micro-averaged F1-score.
Evaluates Named Entity Recognition (NER) capabilities across 11 Indic languages. It probes a model's ability to identify and classify PERSON, LOCATION, and ORGANIZATION entities in low-resource and multilingual settings using projection-based and fine-tuned approaches. Use when the user wants to benchmark on Naamapadam, or asks about evaluating this task. Reports F1.
Evaluates real-time anomaly detection algorithms on streaming time-series data, measuring their ability to detect natural and synthetic anomalies while penalizing false alarms and delayed detections. Use when the user wants to benchmark on NAB 1.0, or asks about evaluating this task. Reports NAB Score.
Evaluates unsupervised anomaly detection algorithms on streaming time-series data. It probes the model's ability to identify point, contextual, and collective anomalies in highly imbalanced real-world and synthetic datasets without labeled training data. Use when the user wants to benchmark on Numenta Anomaly Benchmark, Yahoo Anomaly Dataset, or asks about evaluating this task. Reports F-measure.
Evaluates a model's ability to predict discrete neural audio codec parameters (quantizers, sampling rate, bits per second) from audio samples, enabling fine-grained source attribution of AI-generated speech. The protocol frames open-set attribution as a multi-task regression problem rather than binary classification, requiring the model to generalize across both seen and unseen codec configurations. Use when the user wants to benchmark on ST-Codecfake, CodecFake, or asks about evaluating this...
Evaluates speech-to-text and speech-to-speech translation capabilities across low-resource Nigerian languages (Hausa, Igbo, Yorùbá, Nigerian Pidgin) and English. It specifically probes how well cascaded, end-to-end, and AudioLLM architectures handle multi-accent variations and bidirectional translation directions. Use when the user wants to benchmark on NaijaS2ST, or asks about evaluating this task. Reports SSA-COMET.
Evaluates how pre-training data exposure and external context influence closed-book and open-book question answering accuracy. It probes the model's reliance on parametric knowledge versus retrieved evidence, and measures the impact of answer frequency and distractors. Use when the user wants to benchmark on Natural Questions, SQuAD, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates session-based recommendation models by predicting the next item a user will click based on their sequential interaction history within a session. It probes the model's ability to capture both sequential behavior and session-level intent/purpose. Use when the user wants to benchmark on YOOCHOOSE 1/64, YOOCHOOSE 1/4, DIGINETICA, or asks about evaluating this task. Reports Recall@20.
This benchmark evaluates a model's ability to perform abstractive and extractive summarization on long-form narrative texts (movie/TV plot descriptions). It probes the model's capacity to capture event causality, character motivations, temporal dynamics, and overall narrative coherence while maintaining faithfulness to the source document. Use when the user wants to benchmark on NarraSum, or asks about evaluating this task. Reports ROUGE F1.
Evaluates English long-document retrieval on complex, narrative-style questions, probing deep comprehension and information extraction from lengthy texts. Use when the user wants to benchmark on NarrativeQA, or asks about evaluating this task. Reports nDCG@10.
Evaluates the ability of deep learning and geometric deep learning models to forecast county-level COVID-19 hospitalizations using satellite-derived atmospheric variables (AOD, temperature, humidity) alongside baseline features. It probes spatio-temporal forecasting capabilities and the conditional predictive utility of environmental risk factors on disease severity. Use when the user wants to benchmark on NASAdat, or asks about evaluating this task. Reports RMSE.
Evaluates the generation quality and inference efficiency of structuredly pruned encoder-decoder language models across abstractive QA, summarization, classification, and instruction-following tasks. Use when the user wants to benchmark on TweetQA, XSum, SAMSum, CNN/DailyMail, GLUE/SuperGLUE (RTE, BoolQ, CB), Databricks-dolly-15k, Self-Instruct, Vicuna Evaluation, or asks about evaluating this task. Reports ROUGE-L.
Compute NathanFradet/ece via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of NathanFradet/ece.
Compute NathanFradet/levenshtein via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of NathanFradet/levenshtein.
Evaluates a model's ability to generalize to unseen NLP tasks by leveraging crowdsourced natural language instructions alongside training data. It measures how well instruction-based learning transfers across different task categories, datasets, and individual tasks compared to data-only training. Use when the user wants to benchmark on Natural Instructions, or asks about evaluating this task. Reports ROUGE-L.
Evaluates the zero-shot reasoning capabilities of models trained via knowledge distillation or self-training on the NaturalReasoning dataset. It measures performance across diverse mathematics and science benchmarks to assess scaling efficiency and generalization. Use when the user wants to benchmark on MATH, GPQA, GPQA-Diamond, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
Evaluates the perceptual quality and generation speed of an end-to-end text-to-speech system. It measures how closely synthesized speech matches human recordings and outperforms prior cascaded or flow-based TTS baselines. Use when the user wants to benchmark on LJSpeech, or asks about evaluating this task. Reports CMOS.
Evaluates the ability of voice conversion models to preserve speaker identity, intelligibility, and emotional expression when converting spontaneous, in-the-wild podcast speech. It benchmarks both standard and emotion-aware conversion across multiple architectures and data scales. Use when the user wants to benchmark on NaturalVoices, ESD, or asks about evaluating this task. Reports WER.
Evaluates large multimodal models on underwater scene understanding across eight tasks, including coarse/fine classification, image/region captioning, grounding, detection, VQA, and object counting. It probes the model's robustness to severe underwater image degradation (light scattering, absorption, color casts) and its ability to generalize to unseen underwater domains. Use when the user wants to benchmark on NautData, IOCfish5k, MarineInst20M, or asks about evaluating this task. Reports ac...
Evaluates a voice cloning system's ability to generate high-quality, speaker-similar speech using minimal untranscribed or transcribed target speech. It probes both text-to-speech (TTS) and voice conversion (VC) capabilities, focusing on naturalness, speaker similarity, and accent preservation across native and non-native speakers. Use when the user wants to benchmark on VCC2018 SPOKE task, VCTK & EMIME, or asks about evaluating this task. Reports MOS.
Evaluates closed-loop end-to-end autonomous driving performance in safety-critical and diverse real-world scenarios. It measures collision avoidance, rule compliance, progress, and comfort under reactive traffic conditions. Use when the user wants to benchmark on navhard, navtest, or asks about evaluating this task. Reports EPDMS.
Evaluates end-to-end autonomous driving planners on trajectory prediction and safety-critical behaviors. It measures compliance with traffic rules, drivable area boundaries, collision avoidance, and driving comfort over short-horizon scenarios. Use when the user wants to benchmark on NAVSIM, or asks about evaluating this task. Reports PDM Score (PDMS).
Evaluates time series anomaly detection models across unsupervised, semi-supervised, and supervised settings. It probes the model's ability to detect point and segment anomalies using window-based contextual representations and outlier exposure techniques. Use when the user wants to benchmark on SMAP, MSL, SWaT, SMD, Yahoo, KPI, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to predict implicit user-item interactions and rank relevant items for recommendation. It probes non-linear collaborative filtering capabilities on sparse, implicit feedback datasets by measuring whether the true interacted item appears near the top of a ranked list. Use when the user wants to benchmark on MovieLens, Pinterest, or asks about evaluating this task. Reports HR@10, NDCG@10.
Probes a model's ability to perform fine-grained information extraction from scholarly NLP papers. It specifically tests the identification of contribution sentences, the extraction of scientific terms and entities, and the structuring of these elements into RDF-style triples organized under 12 predefined information units. Use when the user wants to benchmark on NLPContributionGraph, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to retrieve relevant documents from a large corpus given a natural language query. It measures ranking quality and recall at various cutoffs to assess end-to-end document retrieval performance. Use when the user wants to benchmark on NQ320k, TriviaQA, or asks about evaluating this task. Reports Recall@1.
This evaluation probes the ability of a unified neural cellular automaton substrate to simultaneously evolve robot morphology and control policies for navigation and manipulation tasks. It measures how well decentralized, local interactions enable robots to chase light, navigate obstacles, and carry objects in custom simulation environments. Use when the user wants to benchmark on NCRS Custom Simulation Environments (LC, LCO, CBT), or asks about evaluating this task. Reports fitness score.
Compute NCSOFT/harim_plus via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of NCSOFT/harim_plus.
Compute the ndcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute ndcg_score, or asks how to score with ndcg_score.