
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates vision-language models on five assistive technology tasks for people with visual impairments: panoptic segmentation, depth estimation, optical character recognition, image captioning, and visual question answering. It probes the model's ability to unify multiple visual understanding and generation tasks within a single parameter set using task-specific prompts. Use when the user wants to benchmark on ADE-150, NYU-V2, OCR (IC13, IC15, SVT, IIIT5K, SVTP, CUTE), VizWiz_Cap, VizWiz_VQA,...
Evaluates automatic speech recognition (ASR) and call sign detection (CSD) capabilities on real-world air traffic control audio. It probes the model's ability to transcribe accented, noisy pilot and controller speech and accurately identify aircraft call signs in domain-specific phraseology. Use when the user wants to benchmark on Airbus ATC Speech Recognition 2018 Challenge Corpus, or asks about evaluating this task. Reports WER.
This benchmark evaluates the robustness of self-supervised speech recognition models under domain shift in air traffic control communications. It probes few-shot fine-tuning capabilities, sensitivity to audio quality and accents, and potential gender bias in transcription performance. Use when the user wants to benchmark on NATS, ISAVIA, LiveATC-Test, ATCO2-Test, LDC-ATCC, UWB-ATCC, ATCOSIM, or asks about evaluating this task. Reports WER.
Evaluates a drone-based suspect-and-investigate system for detecting and classifying illegally parked cars, moving cars, and legally parked cars from aerial imagery. Use when the user wants to benchmark on ATG-PVD, or asks about evaluating this task. Reports mAP.
Evaluates performance-based anomaly detection methods to identify athletes with confirmed anti-doping rule violations. It measures how well different algorithms surface sanctioned athletes while balancing precision and recall, accounting for environmental factors like wind and altitude. Use when the user wants to benchmark on 100 m Sprint, or asks about evaluating this task. Reports F1 score.
This benchmark evaluates human motion prediction algorithms by testing their ability to forecast future trajectories given past observations and environmental context. It systematically probes robustness to perception noise, generalization across different social and cultural environments, and performance under varying observation and prediction horizons. Use when the user wants to benchmark on ETH, ATC, THÖR, or asks about evaluating this task. Reports ADE.
Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capabilities in a low-resource setting. Use when the user wants to benchmark on DarijaMMLU, DarijaHellaSwag, Belebele_Ary, DarijaBench, DarijaAlpacaEval, or asks about evaluating this task. Reports Accu...
This benchmark probes frontier scientific reasoning across multiple disciplines (e.g., physics, chemistry, biology, computer science, mathematics) using original, multi-step problems. It evaluates a model's ability to generate complex, open-ended, LaTeX-formatted answers and assesses both solution accuracy and inference stability across multiple sampling runs. Use when the user wants to benchmark on ATLAS, or asks about evaluating this task. Reports Accuracy.
Evaluates long-term personalized referential memory QA by testing a model's ability to retrieve and reason over multi-source, multimodal personal data spanning years. It probes conflict-aware aggregation, temporal-visual grounding, and accurate reference resolution across different question types. Use when the user wants to benchmark on ATM-Bench, or asks about evaluating this task. Reports QS.
Evaluates the ability of machine learning models to reconstruct missing multivariate atmospheric data across time and altitude. It probes spatiotemporal continuity, physical gradient preservation, and performance under varying gap lengths (short, medium, long). Use when the user wants to benchmark on SD-WACCM-X synthetic atmospheric data, or asks about evaluating this task. Reports Pearson correlation R.
Evaluates neural network parameterizations for predicting subgrid atmospheric processes (e.g., microphysical tendencies, momentum fluxes) using single-column versus non-local (3x3 grid) inputs. It probes the model's ability to capture mesoscale convective dynamics and frontal systems, and assesses performance across different atmospheric stability regimes. Use when the user wants to benchmark on SAM (Simple Atmospheric Model) simulation, or asks about evaluating this task. Reports R^2.
Evaluates SAR automatic target recognition (ATR) capabilities under realistic, wild conditions. It probes fine-grained vehicle classification and detection robustness across varying imaging geometries, scene complexities, and domain shifts (SOC vs EOC settings). Use when the user wants to benchmark on ATRBench, or asks about evaluating this task. Reports overall accuracy (%).
Evaluates Amur tiger re-identification in the wild by measuring how well models can match tiger identities across different camera views and detection/pose conditions. It probes robustness to non-rigid body deformation, extreme pose variation, and domain shifts between controlled (plain) and uncontrolled (wild) environments. Use when the user wants to benchmark on ATRW, or asks about evaluating this task. Reports mAP.
This benchmark probes the robustness of multimodal large language models (MLLMs) to maliciously manipulated chart visualizations. It measures how well models answer chart-based questions when presented with data-consistent but misleading charts compared to correct charts, and quantifies the rate at which models switch from correct to incorrect answers due to visual misleaders. Use when the user wants to benchmark on AttackViz, or asks about evaluating this task. Reports relaxed accuracy.
Evaluates a session-based recommendation model's ability to predict the next item in a user's browsing session. It probes the model's capacity to capture multi-level user intent and item semantics through attention mechanisms while handling session-specific inductive biases. Use when the user wants to benchmark on Diginetica, Gowalla, Last.fm, or asks about evaluating this task. Reports HR@K, MRR@K.
Evaluates the effectiveness of attention head pruning methods in reducing gender bias in large language models while preserving language model utility. It probes the trade-off between fairness mitigation and general language modeling performance across models of varying sizes. Use when the user wants to benchmark on HolisticBias, WikiText-2, or asks about evaluating this task. Reports HolisticBias.
Predicts whether a pair of drugs interacts based on multi-modal similarity features (targets, pathways, side effects, chemical structure, etc.). Evaluates performance on highly imbalanced drug-drug interaction datasets using precision-recall metrics. Use when the user wants to benchmark on DS1, DS2, DS3 (CYP), DS3 (NCYP), or asks about evaluating this task. Reports AUPR.
Evaluates algorithmic reasoning and out-of-distribution generalization in Transformers by measuring prediction accuracy and attention pattern alignment against ground-truth reference masks on synthetic tasks. Use when the user wants to benchmark on AttentionSpan, or asks about evaluating this task. Reports Accuracy.
Evaluates the feasibility and performance of running standard AI safety benchmarks inside Trusted Execution Environments (TEEs) using quantized models. It probes zero-shot reasoning accuracy, toxicity refusal capabilities, and text generation quality under hardware and cryptographic constraints. Use when the user wants to benchmark on MMLU, ToxicChat, Summarization, or asks about evaluating this task. Reports MMLU Accuracy (%).
Evaluates session-based recommendation models enhanced with the AttrGAU framework on their ability to predict the next item in a user session. It probes robustness to data sparsity and noisy interactions, and measures the model-agnostic performance gain over vanilla backbones. Use when the user wants to benchmark on Dressipi, Diginetica, Retailrocket, or asks about evaluating this task. Reports HR@N, MRR@N.
Evaluates the faithfulness and robustness of perturbation-based image attribution methods by measuring how well generated heatmaps localize objects, predict probability drops upon feature removal, and maintain consistency under hyperparameter variations. Use when the user wants to benchmark on ImageNet, Places365, or asks about evaluating this task. Reports deletion metric.
This evaluation probes a model's ability to perform open-vocabulary semantic segmentation using either direct class names or decomposed attribute descriptions. It specifically tests robustness to textual ambiguity, neologisms, and unnameable categories by measuring pixel-level alignment across standard and novel datasets. Use when the user wants to benchmark on PASCAL-5i, COCO-20i, PASCAL VOC, PASCAL Context, Fantastic Beasts, or asks about evaluating this task. Reports mIoU.
This benchmark probes the ability of deepfake audio detectors to generalise to novel synthesis systems and diverse human voice corpora in open-world settings. It specifically evaluates robustness against domain shifts in both speech synthesis methods and real audio sources, revealing how well models handle unseen acoustic patterns and distribution shifts. Use when the user wants to benchmark on AUDETER, or asks about evaluating this task. Reports Equal Error Rate (EER).
Probes whether audio encoders preserve algebraic consistency when identical sources are added to different base scenes (A-COAT), and whether representations can be accurately reconstructed from discrete attribute-level primitives like timbre, pitch, rate, and amplitude (A-TRE). Use when the user wants to benchmark on Synthetic Audio Scenes, or asks about evaluating this task. Reports A-COAT.
This protocol evaluates machine translation quality by comparing crowd-sourced human judgments of text-only outputs versus multimodal (text + audio) outputs. It probes whether audio-based assessments improve inter-rater consistency and reveal system-level differences through prosodic and expressive features unavailable in text. Use when the user wants to benchmark on WMT German-English, or asks about evaluating this task. Reports standardized score.
Evaluates pretrained audio deepfake detectors across 28 diverse datasets to measure robustness against different manipulation types, generation methods, and real-world conditions like in-the-wild noise and perturbations. The protocol standardizes audio preprocessing and label formats to enable fair cross-dataset comparison and highlights generalization gaps when lab-trained models face advanced generation techniques. Use when the user wants to benchmark on ASVspoof2019_LA, MLAAD-v5, In-the-wi...
This benchmark evaluates the generalization capability of audio deepfake detection models by testing their performance on controlled, studio-recorded spoofing data versus real-world, uncontrolled in-the-wild audio. It probes whether models trained on standard lab benchmarks can robustly distinguish real from synthetic speech in practical deployment scenarios. Use when the user wants to benchmark on ASVspoof 2019 LA, In-the-Wild Data, or asks about evaluating this task. Reports EER.
Evaluates the quality, diversity, and realism of generated audio drum loops. It probes a model's ability to capture spectral-temporal patterns, genre characteristics, and seamless looping properties in fixed-length music generation. Use when the user wants to benchmark on FreeSound Loop Dataset (FSLD), or asks about evaluating this task. Reports IS.
Evaluates an audio-only model's ability to disentangle and reconstruct individual sound sources from mixed audio using semantic category guidance, without relying on visual cues. Use when the user wants to benchmark on MUSIC, FUSS, MUSDB18, VGG-Sound, or asks about evaluating this task. Reports SDR.
Evaluates the robustness of audio spoof detection models against real-world audio degradation and manipulation attacks (laundering), including reverberation, additive noise, and re-compression. Use when the user wants to benchmark on ASVspoof 2019 LA, ASVspoof Laundered Database, or asks about evaluating this task. Reports EER.
Evaluates an audio-to-image generative model's ability to synthesize semantically aligned images from audio prompts. It measures cross-modal alignment, perceptual image quality, and distributional similarity against ground-truth visuals. Use when the user wants to benchmark on Greatest Hits, Landscapes, Into The Wild (ITW), VEGAS, VGGSound, or asks about evaluating this task. Reports Fréchet Inception Distance (FID).
Evaluates the human-likeness of Chinese text-to-speech systems using a Turing-test-inspired protocol where human listeners classify audio as human, unclear, or machine. It also benchmarks an automatic LLM-based evaluator against human judgments and traditional MOS prediction models to measure alignment and trap-item detection capability. Use when the user wants to benchmark on ATT-Corpus, or asks about evaluating this task. Reports HLS.
Evaluates audio-visual few-shot video classification by measuring how well models generalize to novel classes with limited training examples (1, 5, 10-shot). It reports both generalised few-shot learning (HM) and standard few-shot learning (FSL) accuracy to assess bias towards base classes. Use when the user wants to benchmark on VGGSound-FSL, UCF-FSL, ActivityNet-FSL, or asks about evaluating this task. Reports HM.
Probes an embodied agent's ability to navigate unmapped 3D environments using fused audio-visual observations to locate both static and moving sound sources. It tests generalization to unseen environments and unheard audio distributions under clean and noisy conditions. Use when the user wants to benchmark on Replica, Matterport3D, or asks about evaluating this task. Reports Success rate (SR).
This protocol evaluates an audio-visual model's ability to isolate and separate target instrument sounds from mixed multi-source video audio using visual object cues. It quantifies separation accuracy and artifact suppression across held-out test clips and synthetically mixed pairs. The evaluation also probes generalization to unseen object combinations and visually-guided denoising on real-world videos. Use when the user wants to benchmark on MUSIC, AudioSet-Unlabeled, AudioSet-SingleSource,...
Evaluates zero-shot language-queried audio source separation using natural language captions rather than fixed labels. The benchmark tests the model's ability to separate a target sound described by human-annotated captions from a mixed audio mixture. Use when the user wants to benchmark on AudioCaps, or asks about evaluating this task. Reports SDRi.
Evaluates speech-in-speech-out dialogue systems on their ability to accurately answer spoken queries using external tools, measuring both answer correctness and system latency under streaming versus open-book settings. Use when the user wants to benchmark on AudioCRAG, or asks about evaluating this task. Reports accuracy.
Evaluates long-context audio understanding and inference efficiency across speech, sound, and music domains. It probes temporal dependency modeling, multi-hop reasoning, and memory/token-pruning scalability in Large Audio Language Models. Use when the user wants to benchmark on AudioMarathon, or asks about evaluating this task. Reports F1-score.
Evaluates the robustness of audio watermarking detectors against watermark removal and forgery under no-box, black-box, and white-box threat models. It also measures the perceptual and signal quality of perturbed audio to ensure attacks do not excessively degrade the host signal. Use when the user wants to benchmark on AudioMarkData, LibriSpeech, or asks about evaluating this task. Reports FNR, FPR.
Evaluates audio classification performance on spoken digits and speaker sex using raw waveforms and spectrograms, serving as a benchmark for explainable AI (XAI) methods in the audio domain. Use when the user wants to benchmark on AudioMNIST, or asks about evaluating this task. Reports accuracy.
Evaluates large audio-language models' ability to perceive and reason about spatial motion in binaural audio. It probes whether models can correctly infer motion direction and trajectories from interaural cues, rather than relying on linguistic or spectral heuristics. Use when the user wants to benchmark on AudioMotionBench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to generate audio responses that simultaneously maintain contextual semantic appropriateness and preserve the specific acoustic identity (timbre, prosody) of a target character given reference audio. Use when the user wants to benchmark on AudioRole-Demo, or asks about evaluating this task. Reports Acoustic Personalization (AP).
Evaluates audio safety guardrails on detecting both audio-native risks (e.g., harmful sound events, voice attributes) and semantic content risks (e.g., jailbreaks, policy violations). It measures joint accuracy where both risk types must be correctly classified, alongside end-to-end inference latency. Use when the user wants to benchmark on AudioSafetyBench, Jailbreak-AudioBench, Nemotron-Content-Safety-Audio, Omni-SafetyBench, AdvWave, or asks about evaluating this task. Reports accuracy.
Evaluates zero-shot language-queried audio source separation on a diverse set of environmental and acoustic sound classes. The benchmark tests the model's ability to isolate a target sound described by a text label from a synthetically mixed audio mixture. Use when the user wants to benchmark on AudioSet, or asks about evaluating this task. Reports SDRi.
Evaluates whether large language models can correctly reason about physicochemical mechanisms in gold nanoparticle synthesis using multiple-choice questions. It probes both factual recall and the depth of mechanistic understanding by measuring prediction accuracy and model confidence derived from output logits. Use when the user wants to benchmark on AuNP Synthesis Mechanism Benchmark, or asks about evaluating this task. Reports accuracy.
Compute the AUROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AUROC, or asks how to score with AUROC.
Evaluates the predictability and forecast skill of the Aurora AI weather model across selected extreme weather events, including tropical cyclones, winter freezes, and heatwaves. It probes the model's ability to maintain deterministic track accuracy, temperature amplitude, and spatial pattern fidelity across short-range (1–7 day) to subseasonal (14–21 day) lead times. Use when the user wants to benchmark on Selected Weather Extremes Case Studies, or asks about evaluating this task. Reports Tr...
Evaluates an LLM's ability to predict the correct legal citation for a given query text. It probes the model's capacity to retrieve or generate accurate references from a large Australian legal corpus, testing both retrieval and generation capabilities in a domain-specific setting. Use when the user wants to benchmark on AusLaw Citation Benchmark, or asks about evaluating this task. Reports ACC@1.
Evaluates an LLM's capability to detect and categorize hallucinations in authentic, real-world human-LLM dialogues. It specifically probes whether models can identify input-conflicting, context-conflicting, and fact-conflicting errors in query-response pairs. Use when the user wants to benchmark on AuthenHallu, or asks about evaluating this task. Reports accuracy.
Evaluates an automated program repair system's ability to fix real-world security and logic bugs in OpenID Connect implementations. It measures both the rate of successfully generating correct patches and the semantic quality of those patches compared to human-written developer fixes. Use when the user wants to benchmark on OpenID Connect Bug Dataset, or asks about evaluating this task. Reports fix_accuracy.