
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the accuracy of event-based optical flow estimation models in underwater environments. It probes how well algorithms handle low-texture, turbid, and refractive conditions compared to terrestrial benchmarks. Use when the user wants to benchmark on UEOF, or asks about evaluating this task. Reports AEE.
Evaluates how data preprocessing choices—such as observation selection, flow directionality, and feature engineering—affect the performance of unsupervised anomaly detection models on network traffic. Use when the user wants to benchmark on UGR'16, or asks about evaluating this task. Reports AUC.
This evaluation probes a model's ability to understand and reason about user interface components across multiple modalities (images, text, structural metadata). It tests cross-modal alignment, component retrieval, synchronization detection, and classification tasks relevant to UI design and accessibility. Use when the user wants to benchmark on Rico, or asks about evaluating this task. Reports accuracy.
Evaluates mobile agents' ability to complete long-horizon, dependency-rich tasks on real mobile applications. It specifically probes atomic-to-compositional generalization, testing how well agents handle task concatenation, context transitions, and deep analysis across different app types and languages. Use when the user wants to benchmark on UI-NEXUS, or asks about evaluating this task. Reports Success Rate.
Evaluates Vietnamese language models' machine reading comprehension capabilities, specifically probing their ability to extract correct answer spans and correctly identify when a question cannot be answered from the given context. Use when the user wants to benchmark on UIT-ViQuAD 2.0, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates the impact of data assimilation (DA) using the SPEnKF algorithm on a U-STN12 deep learning model for UK temperature forecasting. It probes the model's ability to integrate global atmospheric data (ERA5 T850) and surface observations (ASOS/ERA5 T2m) over a 120-hour lead time, measuring forecast accuracy degradation or improvement under varying noise levels and assimilation frequencies. Use when the user wants to benchmark on ERA5, ASOS, or asks about evaluating this task. Reports RMSE.
Evaluates the ability of generative models to synthesize realistic 10-meter hourly wind speed maps over the UK, focusing on statistical fidelity, spatial structure, and extreme event intensity distribution. Use when the user wants to benchmark on ERA5, or asks about evaluating this task. Reports FID.
Evaluates the ability of multimodal LLMs to predict binary disease risk from individual-specific clinical data, including tabular features and time-series spirograms. It tests how well serialized text and cross-modal embeddings integrate to produce accurate risk scores for conditions like asthma and diabetes. Use when the user wants to benchmark on UK Biobank, or asks about evaluating this task. Reports AUROC.
Evaluates 3D medical image segmentation models on multi-modal MRI and CT scans across abdominal, brain, and whole-body anatomical domains. Probes the model's ability to produce accurate organ and tumor masks while measuring both volumetric overlap and boundary precision under zero-shot and fine-tuning settings. Use when the user wants to benchmark on AMOS, BTCV, BRATS, UKBOB, or asks about evaluating this task. Reports Dice Score.
Evaluates cross-lingual transfer methods for Ukrainian text classification across toxicity, formality, and natural language inference tasks. It compares translation-based baselines, LLM prompting, and adapter/fine-tuning approaches on both machine-translated and semi-natural Ukrainian test sets. Use when the user wants to benchmark on Ukrainian Toxicity (Translated & Semi-natural), Ukrainian Formality (Translated & Semi-natural), Ukrainian NLI (Translated & Semi-natural), or asks about evalua...
Evaluates the quality of pre-training datasets by measuring the downstream performance of models trained on them. It probes general language understanding, commonsense reasoning, and multilingual capabilities through standard zero-shot benchmarks. Use when the user wants to benchmark on MMLU, ARC-C, ARC-E, CommonSenseQA, HellaSwag, OpenbookQA, PIQA, SIQA, Winogrande, C-Eval, CMMLU, or asks about evaluating this task. Reports Average.
Evaluates a chat model's ability to generate accurate, informative, and correct responses across diverse domains including commonsense, world knowledge, professional knowledge, mathematics, reasoning, and writing. It probes both factual correctness and response quality using automated LLM-based pairwise and independent scoring. Use when the user wants to benchmark on UltraChat Evaluation Set, or asks about evaluating this task. Reports ChatGPT scoring.
Evaluates audio foundation models across understanding, generation, and codec capabilities. It probes semantic accuracy, timbre fidelity, acoustic quality, and multilingual speech comprehension using a unified taxonomy and standardized benchmarks. Use when the user wants to benchmark on SpeechCMMLU, SpeechHSK, LibriSpeech, AISHELL-1, or asks about evaluating this task. Reports WER.
Evaluates text-to-image diffusion models on ultra-high-resolution generation, probing semantic alignment with prompts, fine-grained texture preservation, and overall perceptual quality at resolutions ≥4096px. Use when the user wants to benchmark on UltraHR-eval4K, Aesthetic-Eval@4096, or asks about evaluating this task. Reports FID.
Evaluates multilingual LLMs on chat, math reasoning, and code generation across five languages (English, Chinese, Spanish, Russian, French) to measure the effectiveness of knowledge-enhanced supervised fine-tuning. Use when the user wants to benchmark on OMGEval, MGSM, Multilingual HumanEval, or asks about evaluating this task. Reports OMGEval score.
Evaluates spoken dialogue models' ability to follow fine-grained speech style instructions (emotion, speed, volume, accent, language, composite) while maintaining general conversational competence. It measures both subjective audio quality/naturalness and objective content/emotion alignment against ground-truth style specifications. Use when the user wants to benchmark on UltraVoice Test Set, URO-Bench, or asks about evaluating this task. Reports MOS, IFR.
This protocol evaluates the reliability of automatically generated relevance judgments (via the UMBRELA tool) compared to human assessments across different workflow conditions. It measures how well LLM-generated qrels align with human qrels in ranking retrieval systems using standard IR metrics and rank correlation. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports Kendall's τ.
Evaluates the cross-modality generalization and robustness of 3D medical segmentation foundation models by testing their ability to segment 13 whole-body organs in functional (PET) versus structural (CT/MRI) imaging using intrinsically paired intra-subject scans. Use when the user wants to benchmark on UMD Benchmark, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
Evaluates the ability of deep learning models to perform fine-grained named entity recognition for UMLS semantic types in biomedical text. It probes how well models handle class imbalance, contextual ambiguity, and domain shift between clinical notes and biomedical abstracts. Use when the user wants to benchmark on i2b2 2010, MedMentions(full), MedMentions(st21pv), or asks about evaluating this task. Reports F1.
Evaluates zero-shot and supervised machine translation quality across multiple language pairs using a multilingual encoder-decoder architecture. It probes the model's ability to translate between unseen language pairs (e.g., Spanish-French) using only monolingual data and reinforcement learning, without parallel training data for the target pair. Use when the user wants to benchmark on United Nations Parallel Corpus (UN corpus), or asks about evaluating this task. Reports BLEU.
Evaluates a model's ability to generate unbiased scene graphs by predicting pairwise relationships between objects in images. It specifically probes robustness to long-tailed predicate distributions by measuring per-class recall averaged across all predicate classes, rather than relying on global recall which favors head classes. Use when the user wants to benchmark on VG150, GQA200, or asks about evaluating this task. Reports mR@K.
Evaluates the robustness, calibration, and selective classification capability of uncertainty estimation methods (Deep Ensembles, MC Dropout, SVI, TTA) on histopathological whole slide images under domain shift and label noise. Use when the user wants to benchmark on Camelyon17, TCGA, or asks about evaluating this task. Reports AUARC.
Benchmarks the robustness of uncertainty estimation methods against label outliers and distribution shifts. It evaluates whether predicted prediction intervals and uncertainty quantifications maintain calibration and accuracy when training data is contaminated with noise or adversarial perturbations. Use when the user wants to benchmark on Synthetic 1D regression dataset, Real-world regression datasets, NYU-Depth-v2, or asks about evaluating this task. Reports Interval score.
Evaluates a CNN's ability to predict stellar atmospheric parameters and chemical abundances from low-resolution spectra, measuring both internal consistency across model runs and agreement with established spectroscopic pipeline measurements. Use when the user has predictions and gold and needs to compute Uncertainty.
Evaluates a robot's ability to infer human goals and execute household tasks from noisy, accented, or mispronounced spoken instructions. It probes robust speech perception, joint planning, and Theory of Mind in embodied human-robot collaboration under mixed-observability conditions. Use when the user wants to benchmark on UnclearInstruct, or asks about evaluating this task. Reports Accuracy.
This protocol evaluates unbiased learning-to-rank models on their ability to correct position bias and propensity overestimation using implicit click feedback. It probes ranking quality under both dynamic online and static offline logging policies by comparing predicted rankings against ground truth relevance. Use when the user wants to benchmark on Yahoo! LETOR, Istella-S, or asks about evaluating this task. Reports NDCG@K.
Evaluates code generation on novel, synthetic programming problems designed to avoid training data contamination. It assesses the model's ability to understand and implement complex logic beyond simple unit test passing. Use when the user wants to benchmark on Unconventional Problems, or asks about evaluating this task. Reports LLM graded Understanding score.
Evaluates pixel-level biomedical image segmentation capability using convolutional networks. Probes the model's ability to precisely delineate cellular structures and membranes in electron and light microscopy images with limited training data. Use when the user wants to benchmark on EM segmentation challenge (ISBI 2012), PhC-U373, DIC-HeLa, or asks about evaluating this task. Reports warping error, IOU.
Evaluates retrieval and end-to-end generation performance of multimodal RAG systems on real-world PDF documents. It probes the ability of text-only, image-only, and multimodal (text-image fusion/joint) retrieval paradigms to locate relevant evidence and generate faithful, complete answers to cross-modality questions. Use when the user wants to benchmark on UNIDOC-BENCH, or asks about evaluating this task. Reports Precision@10, Recall@10, Faithfulness, Completeness.
Evaluates an instance segmentation model's ability to accurately delineate and separate individual microstructural objects in high-resolution electron micrographs, particularly under varying instance densities. Use when the user wants to benchmark on UniEM-3M, or asks about evaluating this task. Reports mAP@0.5.
Evaluates natural language generation models across multiple quality dimensions (e.g., coherence, fluency, consistency, relevance) by reframing assessment as a Boolean QA task. Measures how well automated scores align with human judgments using correlation metrics. Use when the user wants to benchmark on SummEval, Topical-Chat, SFRES, SFHOT, QAGS, or asks about evaluating this task. Reports Spearman correlation.
Evaluates large language models' factual correctness by dynamically generating responses to factual questions and measuring how well hallucination detection and fact verification methods can predict the ground-truth factuality label of those responses. It probes a model's susceptibility to hallucination and the effectiveness of external evidence retrieval in verifying generated claims. Use when the user wants to benchmark on TriviaQA, NQ-Open, PopQA, 2WikiMultihopQA, HotpotQA, or asks about e...
Evaluates medical vision-language models across a broad capability surface including visual diagnosis, medical imaging, clinical reasoning, text-based QA, report generation, and instruction following. It emphasizes protocol reproducibility, deployment relevance (safety, consistency, faithfulness), and robustness to real-world clinical imagery and OCR conditions. Use when the user wants to benchmark on Public Med-VLM Benchmarks (30+ subsets), Inhouse VQA, Inhouse OCR, Inhouse Caption, or asks ...
Evaluates a unified discrete diffusion framework's capability to jointly generate and reason over image-text pairs. It probes unconditional and conditional generation quality, the effectiveness of classifier-free guidance, training and inference efficiency, and cross-modal retrieval and reasoning performance. Use when the user wants to benchmark on DataComp1B, CC12M, MS-COCO30k, Flickr, Winoground, or asks about evaluating this task. Reports FID.
Evaluates multimodal large language models on perceptual-level image understanding across three domains: Image Aesthetics & Art (IAA), Image Quality Assessment (IQA), and Image Structure & Texture Assessment (ISTA). It probes both continuous visual rating (VR) and discrete visual question answering (VQA) capabilities. Use when the user wants to benchmark on UniPercept-Bench, or asks about evaluating this task. Reports Acc..
Evaluates a unified text-to-audio model's ability to generate speech, music, and sound effects from natural language instructions without reference audio. It probes instruction-following fidelity, acoustic quality, structural coherence, and the positive transfer effects of multi-modal joint training. Use when the user wants to benchmark on UniSonate Unified Corpus, Seed-TTS test set, SongEval benchmark, or asks about evaluating this task. Reports WER, SongEval.
Evaluates multimodal voice generation and conversion capabilities across face-driven, text-driven, and attribute-based tasks. Probes the model's ability to align facial, textual, and attribute descriptions with target speech while preserving speaker identity, content clarity, and naturalness. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports MOS-Match.
Evaluates text summarization models across multiple dimensions including faithfulness, completeness, conciseness, domain stability, and abstractiveness. It tests how well summarizers handle diverse input contexts (domains, dialogue vs. non-dialogue, short vs. long texts) and the impact of PII redaction on hallucination. Use when the user wants to benchmark on UniSumEval, or asks about evaluating this task. Reports faithfulness.
Evaluates text-to-SQL models on compositional generalization, out-of-domain robustness, and schema-question alignment across 18 diverse datasets and 12 domains. It probes the model's ability to handle long-form query decomposition and cross-domain SQL pattern diversity. Use when the user wants to benchmark on UNITE, or asks about evaluating this task. Reports accuracy.
This evaluation measures the transcription accuracy of a fine-tuned automatic speech recognition model across four diverse speech benchmarks. It specifically probes the model's robustness to different speaking styles, accents, and linguistic contexts after applying a noise reduction step and a BART-based semantic correction pipeline. Use when the user wants to benchmark on LibriSpeech, Europarl-ASR, TED-LIUM, FLEURS, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates a video agent's capabilities across generation, understanding, editing, and segmentation tasks, while probing its agentic planning and memory mechanisms. It measures how well a unified agent architecture handles long-horizon, multi-step video workflows compared to monolithic baselines. Use when the user wants to benchmark on UniVA-Bench, or asks about evaluating this task. Reports MLLM Judge.
Evaluates video foundation models across six core tasks (understanding, generation, editing, reconstruction) by scoring their outputs on eight cinematic and semantic dimensions, including subject consistency, action dynamics, camera movement, and lighting. Use when the user wants to benchmark on UniVBench, or asks about evaluating this task. Reports UniV-Eval.
This benchmark probes an LLM agent's ability to perform spatial-temporal reasoning and generate executable code for Earth Observation tasks. It evaluates whether models can correctly answer yes/no questions derived from scientific articles by leveraging remote sensing data via Google Earth Engine. Use when the user wants to benchmark on UnivEARTH, or asks about evaluating this task. Reports accuracy.
Evaluates the cross-lingual and cross-domain generalization of decoder-based language models finetuned via contrastive learning on English data. Probes the model's ability to generate unified embeddings for natural language and code retrieval, semantic textual similarity, and intent classification across diverse languages and domains. Use when the user wants to benchmark on MTEB, CodeSearchNet, Multi-CPR, MASSIVE, STS-17 & STS-22, MIRACL, BUCC, or asks about evaluating this task. Reports Spea...
Evaluates multilingual named entity recognition (NER) capabilities across 22 languages and 30 datasets, probing both in-language performance and cross-lingual transfer. It also benchmarks large language models as annotators against human inter-annotator agreement to assess guideline adherence and annotation quality. Use when the user wants to benchmark on UNER v2, or asks about evaluating this task. Reports micro F1.
Evaluates a model's ability to recognize human actions from heterogeneous skeleton data with varying joint counts and topologies. It probes cross-domain generalization, zero-shot/few-shot transfer, and robustness to structural discrepancies between sensing modalities. Use when the user wants to benchmark on NTU-60, HumanML3D, NW-UCLA, NTU-120, or asks about evaluating this task. Reports Accuracy.
Compute the UniversalImageQualityIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute UniversalImageQualityIndex, or asks how to score with UniversalImageQualityIndex.
Evaluates scientific document representation models on multilingual abstracts by measuring tokenization coverage, language modeling perplexity, and embedding quality relative to citation networks. It probes whether models can meaningfully process non-Latin scripts and low-resource languages without degrading to English-only or graph-based heuristics. Use when the user has predictions and gold and needs to compute unknown_token_rate.
Evaluates the ability of recommender systems to efficiently remove specific user interactions or sensitive items (unlearning) while preserving recommendation utility. It probes real-world operational constraints, including handling sequential small-batch deletion requests, domain-specific triggers, and low-latency execution across collaborative filtering, session-based, and next-basket recommendation tasks. Use when the user wants to benchmark on TaFeng, Dunnhumby, Instacart, RSC15, DIGI, NOW...
Compute unnati/kendall_tau_distance via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of unnati/kendall_tau_distance.