
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark meta-evaluates objective music evaluation (OE) metrics by measuring how well their similarity scores and classification outputs align with human subjective judgments. It probes whether automated algorithms can reliably capture human perception of musical quality and distinguish human-composed from AI-generated music across diverse genres and generative models. Use when the user wants to benchmark on Armor, or asks about evaluating this task. Reports correlation coefficient.
Evaluates the effectiveness of semi-structured 2:4 pruning methods on large language models by measuring downstream task accuracy and language modeling perplexity. It specifically tests whether adaptive matrix factorization can preserve model capabilities better than direct weight removal while maintaining inference efficiency. Use when the user wants to benchmark on MMLU, GSM8K, BBH, GPQA, ARC-C, WinoGrande, HellaSwag, Wikitext2, C4, or asks about evaluating this task. Reports Task Accuracy ...
Evaluates a multimodal reward model's ability to judge response quality, detect hallucinations, and follow instructions across text and image inputs. It also probes the model's capacity for agentic tool use, specifically its ability to autonomously invoke visual tools to verify claims and perform fine-grained visual reasoning. Use when the user wants to benchmark on ARMBench-VL, VL-RewardBench, RewardBench-2, V* Bench, HRBench-4K, HRBench-8K, MMERealWorld, or asks about evaluating this task. ...
This benchmark evaluates multimodal large language models on their ability to perform cross-modal audio reasoning. It requires models to integrate multiple audio cues (e.g., speech, environmental sounds, speaker identity) and apply logical inference to answer questions, rather than just performing isolated audio tasks like transcription or classification. Use when the user wants to benchmark on ART (Audio Reasoning Tasks), or asks about evaluating this task. Reports Absolute accuracy.
Probes medical AI agents' ability to perform action-based reasoning on synthetic EHR tasks, specifically targeting data retrieval, temporal aggregation, and threshold-based conditional logic. It measures how well models handle clinical failure modes like missing data, numerical aggregation, and constraint chaining. Use when the user wants to benchmark on ART, or asks about evaluating this task. Reports Success Rate (SR; exact match).
Evaluates the safety vulnerabilities of text-to-image models by measuring how often benign, safe prompts trigger the generation of toxic or unsafe images. It also assesses the diversity and safety of the generated red-teaming prompts themselves. Use when the user wants to benchmark on MSCOCO, or asks about evaluating this task. Reports success ratio under safe prompts (%).
Evaluates the quality and diversity of synthetic images generated by models across ten distinct artistic styles. It probes a model's ability to capture class-conditional and unconditional data distributions while measuring trade-offs between sample fidelity and variety. Use when the user wants to benchmark on ArtBench-10, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).
Evaluates the ability of segmentation models (CNNs, Transformers, diffusion models, and vision foundation models) to detect and classify diverse damage types on analogue media across different material and content categories. It probes cross-media generalization using a leave-one-out protocol and tests the effectiveness of zero-shot, supervised, and text-guided prompting strategies for pixel-level damage localization. Use when the user wants to benchmark on ARTeFACT, or asks about evaluating ...
Compute arthurvqin/pr_auc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of arthurvqin/pr_auc.
Evaluates vision-language models on their ability to detect, spatially localize, and explain visual artifacts in AI-generated images. It probes the model's capacity for fine-grained visual reasoning and artifact-aware grounding beyond standard natural image understanding. Use when the user wants to benchmark on ArtiBench, LOKI, or asks about evaluating this task. Reports accuracy, mIoU, ROUGE.
Evaluates the ability to detect AI-generated music by identifying irreversible residual artifacts from neural audio codecs. It probes robustness across diverse generators, lossy compression codecs, and adversarial source-separation attacks, while measuring false-positive rates on real-world music. Use when the user wants to benchmark on ArtifactBench v1, or asks about evaluating this task. Reports F1.
This evaluation probes an LLM's ability to perform complex mathematical reasoning and multi-turn function calling by autonomously deciding when and how to invoke external tools. It measures the model's capacity for outcome-based agentic reasoning, including state tracking, error recovery, and precise final answer generation without step-level supervision. Use when the user wants to benchmark on MATH-500, AIME, AMC, Olympiad Bench, BFCL v3, τ-bench, or asks about evaluating this task. Reports ...
This benchmark evaluates the factual accuracy and contextual reasoning capabilities of LLMs in music question answering. It specifically probes how well models can retrieve and utilize artist-centric knowledge from a domain-specific database versus relying on parametric memory, comparing zero-shot, RAG, and reranked retrieval strategies. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal retrieval-augmented generation pipeline for deep artwork understanding. It probes the system's ability to retrieve relevant art-historical context from a large corpus, classify artwork attributes (style, genre, artist), and generate grounded, interpretable captions/explanations from image input alone. Use when the user wants to benchmark on WikiFragments, WikiArt/ArtGraph, ArtPedia, SemArt v2.0, PaintingForm, or asks about evaluating this task. Reports NDCG@5, Top-1 Acc...
Evaluates the visual realism and physical fidelity of articulated digital assets for robot learning. It measures geometric detail, reconstruction quality, visual feature alignment with real-world data, and joint motion accuracy under external forces. Use when the user wants to benchmark on ArtVIP, or asks about evaluating this task. Reports joint displacement discrepancy.
Evaluates large vision-language models' ability to comprehend and generate text for scientific figures. It probes capabilities in single and multi-figure captioning, contextualized captioning using in-context examples, and inferring paper titles from figure-caption sequences. Use when the user wants to benchmark on ArXivCap, or asks about evaluating this task. Reports BLEU-2.
Evaluates automatic speech recognition (ASR) models on spontaneous Mandarin-English code-switching in multi-turn conversations. It probes the model's ability to accurately transcribe mixed-language speech under realistic, unscripted conditions with diverse speaker backgrounds. Use when the user wants to benchmark on ASCEND, or asks about evaluating this task. Reports MER.
This benchmark evaluates the clinical reasoning, perception, and diagnostic capabilities of multi-modal large language models across 15 medical specialties. It probes tasks ranging from anatomical and attribute perception to disease identification, staging, treatment planning, and medical report generation. Use when the user wants to benchmark on Asclepius, or asks about evaluating this task. Reports accuracy.
Evaluates a unified audio-visual model's ability to detect which speaker is actively speaking in multi-person video scenes and to enhance speech signals by removing background noise and interference. Use when the user wants to benchmark on AVA-ActiveSpeaker, LRS2, TalkSet, Columbia, MUSAN, or asks about evaluating this task. Reports mAP.
This benchmark evaluates visually grounded interactive planning by testing an agent's ability to dynamically adapt action sequences based on real-time visual observations. It isolates plan adaptation from navigation and low-level manipulation, measuring how well models track environmental state and revise plans under minimal or absent corrective feedback. Use when the user wants to benchmark on AsgardBench, or asks about evaluating this task. Reports success_rate.
This benchmark evaluates the semantic safety and ethical reasoning of vision-language models in robotics. It probes whether models can correctly identify desirable versus undesirable actions across multimodal scenes, real-world injury scenarios, and hypothetical ethical dilemmas. Use when the user wants to benchmark on ASIMOV, or asks about evaluating this task. Reports classification accuracy.
Evaluates LLMs' ability to detect intent deficiencies or overconfidence in user queries and request targeted clarification during multi-turn interactive QA. It measures how well models balance asking clarifying questions versus providing final answers, using rubric-based checkpoints to score clarification quality and final answer accuracy. Use when the user wants to benchmark on AskBench, HealthBench, or asks about evaluating this task. Reports single-turn accuracy.
Evaluates the ability of machine learning classifiers to detect network intrusions and adversarial obfuscations using aggregated bidirectional TCP flow features. It probes whether models can distinguish legitimate traffic from direct and obfuscated attacks without relying on packet payloads. Use when the user wants to benchmark on ASNM-CDX-2009, or asks about evaluating this task. Reports F1-measure.
Evaluates classifier resistance to non-payload-based obfuscation (NPBO) techniques like TCP reordering and retransmissions. It tests detection performance when classifiers are trained without vs. with knowledge of obfuscated attacks. Use when the user wants to benchmark on ASNM-NPBO, or asks about evaluating this task. Reports F1-measure.
Evaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection. Use when the user wants to benchmark on ASNM-TUN, or asks about evaluating this task. Reports F1-measure.
Evaluates a model's ability to identify aspect terms (targets), their categories, sentiment polarity, and exact character positions within restaurant review sentences in Turkish. Use when the user wants to benchmark on SemEval 2016 Turkish Restaurant Reviews, SemEval 2016 English-Translated Restaurant Reviews, or asks about evaluating this task. Reports F1 score.
Binary audio classification to detect the presence of pedestrians in urban environments. It probes a model's ability to distinguish pedestrian activity from background noise under varying spatial radii and pedestrian count thresholds. Use when the user wants to benchmark on ASPED, or asks about evaluating this task. Reports macro-average recall.
Evaluates the perceptual quality of audio-visual speech enhancement models in real-world noisy environments. It probes how well models generalize to natural reverberation, multi-source background noise, and speaker occlusion compared to synthetic training conditions. Use when the user wants to benchmark on ASPIRE, or asks about evaluating this task. Reports MUSHRA.
Evaluates automatic speech recognition (ASR) models on spontaneous speech in Bambara, a low-resource West African language. It probes the models' ability to accurately transcribe audio segments in both a controlled test set and a more heterogeneous benchmark. Use when the user wants to benchmark on Afvoices Test, Nyana Eval, or asks about evaluating this task. Reports WER (%).
This evaluation probes an ASR model's ability to continuously adapt to noisy, rural clinical telephony speech while retaining its baseline performance on standard general-domain speech. It specifically measures the trade-off between target-domain transcription accuracy and catastrophic forgetting of pre-trained linguistic knowledge. Use when the user wants to benchmark on Gram Vaani, Kathbath, or asks about evaluating this task. Reports WER.
Evaluates the training efficiency and final accuracy of end-to-end acoustic models for Automatic Speech Recognition across different CPU-GPU co-processor hardware configurations. It measures how quickly models reach specific accuracy targets (Time-to-Accuracy) and compares final error rates against baseline architectures. Use when the user wants to benchmark on ASR_Dataset, or asks about evaluating this task. Reports TTA.
Evaluates Automatic Speech Recognition systems across diverse acoustic conditions, domains, and linguistic settings. It probes model robustness to read vs. spontaneous speech, clean vs. noisy environments, and demographic bias in multilingual crowdsourced data. Use when the user wants to benchmark on LibriSpeech, Switchboard, TED-LIUM 3, CHiME-6, Common Voice 17.0, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates a label-free regression framework that approximates automatic speech recognition error rates (WER and CER) using multimodal embeddings and predicted transcripts. Probes the tool's robustness across diverse acoustic conditions, domains, and out-of-distribution settings. Use when the user wants to benchmark on LibriSpeech, TED-LIUM, GigaSpeech, SPGISpeech, Common Voice, Earnings22, AMI (IHM), People’s Speech, SLUE-VoXCeleb, Primock57, VoxPopuli Accented, ATCOsim, BERSt, CHiME-6, or as...
Evaluates the robustness of end-to-end automatic speech recognition models to real-world acoustic distortions, including far-field reverberation, mixed sampling rates, low-bitrate codecs, and background noise at varying signal-to-noise ratios. Use when the user wants to benchmark on LibriSpeech, BUT ReverbDB, Hub5 Switchboard & CallHome, AISHELL-2, or asks about evaluating this task. Reports greedy WER (%).
Evaluates the robustness and cross-domain generalization of Automatic Speech Recognition (ASR) models by measuring Word Error Rate (WER) across multiple public and in-house English speech datasets with varying acoustic conditions, sampling rates, and speech types. Use when the user wants to benchmark on LibriSpeech, SwitchBoard & Fisher, WSJ, Common Voice, TED-LIUM v3, Robust Video, CHiME-6, or asks about evaluating this task. Reports WER.
Evaluates multilingual automatic speech recognition and speech translation capabilities across diverse language pairs and domains. It also probes cross-lingual semantic alignment through speech-to-speech retrieval and integration with large language models. Use when the user wants to benchmark on Aishell, LibriSpeech, CoVoSTv2, Fleurs, CommonVoice, MLS, VoxPopuli, or asks about evaluating this task. Reports WER.
Evaluates automatic speech recognition (ASR) performance across English and Croatian by measuring word error rate on multiple held-out test sets. It probes the model's ability to accurately transcribe spoken audio, including handling of punctuation and capitalization. Use when the user wants to benchmark on VoxPopuli, FLEURS, Mozilla Common Voice (MCV12), Hugging Face ASR Leaderboard datasets, or asks about evaluating this task. Reports WER.
Evaluates automatic speech recognition (ASR) models on their robustness and fairness across diverse real-world conditions, including accented speech, rehearsed speech, and spontaneous conversational speech. It specifically probes performance disparities related to speaker accent, gender, and socio-economic background. Use when the user wants to benchmark on ALLSSTAR, NISP, VoxPopuli, Buckeye, CORAAL, or asks about evaluating this task. Reports WER.
Evaluates web agents' ability to perform realistic, time-consuming multi-hop navigation and information retrieval tasks across the open web. It probes planning, memory, dynamic interaction, and robustness against hallucinations and navigation failures. Use when the user wants to benchmark on AssistantBench, or asks about evaluating this task. Reports Acc..
Evaluates automatic speech translation (AST) and automatic speech recognition (ASR) performance on English-French and English-Romanian datasets. It probes the model's ability to convert spoken source audio directly into written target-language translations, and to recognize speech transcripts. Use when the user wants to benchmark on AST LibriSpeech, MuST-C, or asks about evaluating this task. Reports BLEU.
Evaluates a model's ability to jointly extract aspect terms, opinion terms, and their sentiment polarities from text. It probes fine-grained aspect-based sentiment analysis by requiring precise span detection and correct pairing of components within sentences. Use when the user wants to benchmark on 14res, 14lap, 15res, 16res, or asks about evaluating this task. Reports F score.
Evaluates a classifier-based anomaly detection pipeline on simulated astronomical transient light curves. It probes the model's ability to identify rare, out-of-distribution events in real-time without prior exposure to the anomalous classes during training. Use when the user wants to benchmark on Simulated LSST-like transient light curves, or asks about evaluating this task. Reports anomaly score.
Evaluates multimodal large language models' ability to comprehend scientific charts and perform knowledge-intensive reasoning in astronomy. It probes visual understanding, data extraction, numerical calculation, and domain-specific inference. Use when the user wants to benchmark on AstroChart, or asks about evaluating this task. Reports Accuracy (%).
Evaluates a cross-modal foundation model's zero-shot and few-shot regression capabilities on galaxy physical properties (redshift, stellar mass, metallicity, age, sSFR) and its cross-modal similarity search performance, using fixed embeddings without task-specific fine-tuning. Use when the user wants to benchmark on PROVABGS, or asks about evaluating this task. Reports R^2.
Evaluates zero-shot semantic retrieval of astronomical images using natural language queries, specifically probing the model's ability to identify rare galactic phenomena (spirals, mergers, gravitational lenses) without curated training labels. It also measures the impact of VLM-based re-ranking on retrieval precision for rare classes. Use when the user wants to benchmark on HSC survey galaxy images, or asks about evaluating this task. Reports nDCG@10.
Evaluates large language models' ability to act as coding assistants for astronomy-specific scientific workflows. It probes domain-specific API usage, data manipulation, and the generation of research-standard visualizations from natural language queries. Use when the user wants to benchmark on AstroVisBench, or asks about evaluating this task. Reports execution-based evaluation.
This benchmark evaluates the ability of vision-language models to perform multi-modal astronomical reasoning across five distinct observational modalities, including optical imaging, radio interferometry, photometry, light curves, and spectroscopy. It probes whether models can correctly classify celestial objects and interpret physical features, while also testing the impact of prompt guidance and input representation (visual vs. numerical) on classification accuracy and reasoning quality. Us...
Evaluates the robustness of speaker verification systems and anti-spoofing countermeasures against synthesized, voice-converted, and replayed speech attacks. It measures how effectively systems can distinguish genuine speech from spoofed audio and quantifies the real-world impact of spoofing on authentication reliability. Use when the user wants to benchmark on ASVspoof 2019, or asks about evaluating this task. Reports EER, min-tDCF.
Evaluates the robustness of audio deepfake detection models against additive noise and measures how speech enhancement algorithms impact spoof detection accuracy. It probes whether improving perceptual speech quality in noisy environments preserves or degrades the discriminative features needed to distinguish real from spoofed audio. Use when the user wants to benchmark on ASVspoof 2019 LA, or asks about evaluating this task. Reports EER.
This evaluation protocol probes the robustness and generalization of audio deepfake detectors under real-world distortions and across heterogeneous audio types. It specifically tests whether models can maintain reliable binary classification performance when facing unseen generation methods, recording condition shifts, and unknown audio categories without relying on type-specific labels. Use when the user wants to benchmark on AT-ADD Challenge Dataset, or asks about evaluating this task. Repo...