
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates object detection and pose classification capabilities on historical European paintings. Probes a model's ability to recognize culturally heritage-specific entities and human-like poses in artistic contexts rather than natural photographs. Use when the user wants to benchmark on DEArt, or asks about evaluating this task. Reports mAP@0.5.
Evaluates transformer-based models on word-level extractive summarization for policy debate evidence. It measures how well models can identify and extract relevant tokens to form summaries of debate arguments. Use when the user wants to benchmark on DebateSum, or asks about evaluating this task. Reports ROUGE F1.
This benchmark evaluates a model's ability to perform multitask learning across ten diverse natural language processing tasks by framing them as a unified question-answering problem. It probes zero-shot generalization, domain adaptation, and the effectiveness of anti-curriculum training strategies without relying on task-specific modules. Use when the user wants to benchmark on decaNLP, or asks about evaluating this task. Reports decaScore.
Evaluates zero-shot question answering robustness against social biases, specifically testing how models adapt to ambiguous versus unambiguous contexts without relying on internal stereotypical knowledge. Use when the user wants to benchmark on BBQ, or asks about evaluating this task. Reports accuracy.
Probes the model's tendency to exhibit deceptive alignment, including alignment faking in chain-of-thought reasoning, jailbreak success rates, and strategic behavior shifts between evaluation and deployment stages. Use when the user wants to benchmark on DECEPTIONBENCH, StrongReject, JailbreakBench, BeaverTails, HarmfulQA, or asks about evaluating this task. Reports DTR.
This benchmark evaluates a model's ability to decipher undersegmented ancient scripts by aligning unknown character sequences with known language stems. It probes phonological reasoning and unsupervised segmentation capabilities without relying on known language proximity or complete word boundaries. Use when the user wants to benchmark on Gothic, Ugaritic, Iberian, or asks about evaluating this task. Reports P@10.
Evaluates a robot policy's ability to perform bimanual dexterous manipulation tasks under varying levels of tactile dependency. It probes visual-propriocceptive coordination, dynamic object interaction, and contact-rich force control. Use when the user wants to benchmark on DECO-50, or asks about evaluating this task. Reports Success Rate.
Evaluates a split-inference system's ability to perform real-time semantic segmentation on driving video streams while compensating for simulated network communication delays. It probes temporal prediction capabilities and feature fusion under latency constraints. Use when the user wants to benchmark on BDD100K, or asks about evaluating this task. Reports mIoU.
Evaluates the ability of a deep learning model to detect and classify structural bias in heuristic optimization algorithms by analyzing raw performance distributions against a uniform null hypothesis. Use when the user wants to benchmark on BIAS toolbox heuristic pool on $f_0$, or asks about evaluating this task. Reports detection accuracy.
Evaluates the effectiveness of a three-stage neural network compression pipeline (pruning, trained quantization, and Huffman coding) in reducing model storage size while preserving classification accuracy on standard computer vision benchmarks. Use when the user wants to benchmark on MNIST, ImageNet (ILSVRC-2012), or asks about evaluating this task. Reports Top-1 Accuracy.
Evaluates a model's ability to learn optimal dynamic hedging strategies for financial derivatives under discrete trading and varying risk preferences. The protocol simulates market paths using a Heston stochastic volatility model and trains a neural network to minimize a convex risk measure of the terminal hedging error. Performance is assessed out-of-sample against a theoretical benchmark. Use when the user wants to benchmark on Discretized Heston model, or asks about evaluating this task. R...
This benchmark evaluates a model's ability to infer high-resolution (1km×1km, hourly) PM2.5 concentrations across an urban area using sparse mobile and fixed sensor data combined with multi-scale urban features. It probes spatial-temporal prediction capabilities and measures how well the model integrates local, neighboring, and macro-scale regional transport dynamics to improve air quality estimation accuracy. Use when the user wants to benchmark on Beijing PM2.5 Mobile Sensing Dataset, or as...
Evaluates an agent's ability to perform long-horizon, multi-step web research to answer complex factual questions. It probes the model's capacity for iterative search, evidence aggregation, and adaptive reasoning under both reproducible offline constraints and live web environments. Use when the user wants to benchmark on BrowseComp-Plus, BrowseComp, GAIA, xbench-DeepSearch, or asks about evaluating this task. Reports accuracy.
Evaluates end-to-end speech recognition accuracy across diverse acoustic conditions including clean read speech, accented speech, and noisy speech in English and Mandarin. It benchmarks model performance against both automated baselines and human transcribers to measure real-world applicability. Use when the user wants to benchmark on WSJ eval'92, WSJ eval'93, LibriSpeech test-clean, LibriSpeech test-other, VoxForge Accented Speech, CHiME eval clean, CHiME eval real, CHiME eval sim, Baidu int...
This benchmark evaluates the ability of multi-modal embedding classifiers to distinguish real human motion videos from AI-generated ones. It probes semantic consistency detection, robustness to video laundering (resolution/compression), and generalization to unseen generative models. Use when the user wants to benchmark on DeepAction, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict click-through rates for display advertisements by combining raw image pixels with contextual features. It probes the model's capacity to learn high-level visual semantics and complex nonlinear interactions for ranking and probability calibration in a highly imbalanced, real-world advertising setting. Use when the user wants to benchmark on Commercial Display Ad Dataset (2015), or asks about evaluating this task. Reports relative AUC.
Evaluates the emotional expressivity and transferability of a generated multi-turn spoken dialogue dataset by training speech emotion recognition models and measuring their classification performance on held-out and zero-shot test sets. Use when the user wants to benchmark on DeepDialogue (SER subset), RAVDESS, or asks about evaluating this task. Reports accuracy.
Binary classification of protein sequences to determine if they are extracellular matrix (ECM) proteins. It evaluates the model's ability to handle class imbalance and generalize across different species and feature extraction methods. Use when the user wants to benchmark on benchmark dataset, independent dataset, ECMPride dataset, or asks about evaluating this task. Reports balanced accuracy.
Evaluates large vision-language models on fine-grained visual perception, grounding, hallucination mitigation, and multimodal reasoning. It specifically probes the model's ability to autonomously use image zoom-in tools for interleaved visual-linguistic reasoning (iMCoT) to solve high-resolution and complex visual tasks. Use when the user wants to benchmark on V* Bench, HR-Bench, refCOCO / refCOCO+ / refCOCOg / ReasonSeg, POPE, MathVista, MathVerse, or asks about evaluating this task. Reports...
Evaluates the robustness of audio-based biometric authentication systems against deepfake speech synthesis attacks. It measures how easily voice cloning models can bypass speaker verification and how effectively anti-spoofing detectors can distinguish genuine from synthetic speech. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports Bypass Rate.
Evaluates deepfake voice detection models on their ability to generalize from controlled lab synthetic speech to real-world presentation distortions like loudspeaker playback and telephony injection. It measures robustness against realistic signal dynamics and spoofing pipelines that degrade audio quality. Use when the user wants to benchmark on ASVspoof19 LA, ASVspoof21 LA, ASVspoof21 LA-HT, ASVspoof21 DF, ASVspoof5 w/o Enc., In-the-wild, SpoofCeleb, Realworld, or asks about evaluating this ...
Evaluates click-through rate (CTR) prediction models by measuring their ability to correctly rank clicked versus non-clicked instances and output calibrated click probabilities. Use when the user wants to benchmark on Criteo Dataset, Company* Dataset, or asks about evaluating this task. Reports AUC.
Evaluates furniture detection, segmentation, instance retrieval, and set retrieval in indoor scenes. It probes occlusion robustness, fine-grained attribute-based feature learning, and spatial co-occurrence modeling for interior design understanding. Use when the user wants to benchmark on DeepFurniture, or asks about evaluating this task. Reports AP, ACC@K.
Evaluates a unified multimodal model's capabilities in text-to-image generation, image editing, and world-knowledge reasoning. It probes semantic alignment, long-horizon instruction following, fine-grained attribute binding, and precise text rendering across diverse scenarios. Use when the user wants to benchmark on GenEval, DPG-Bench, UniGenBench, WISE, T2I-CoREBench, ImgEdit, GEdit-EN, UniREditBench, RISE, CVTG-2K, or asks about evaluating this task. Reports GenEval.
Binary classification capability to distinguish background noise from gravitational-wave signals (specifically BBH and SGLF classes) in time-series data. It probes the model's ability to generalize to unseen gravitational wave anomalies using deep latent features. Use when the user wants to benchmark on HDR A3D3 gravitational-wave dataset, or asks about evaluating this task. Reports AUC.
Compute the DeepImageStructureAndTextureSimilarity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute DeepImageStructureAndTextureSimilarity, or asks how to score with DeepImageStructureAndTextureSimilarity.
Evaluates a model's ability to perform counterfactual reasoning in biomedical settings by predicting clinical trial outcomes under perturbations of either the outcome measure or the study arm, using similarity-based retrieval to construct counterfactual pairs. Use when the user wants to benchmark on CT open evaluation sample, or asks about evaluating this task. Reports clinical trial outcome prediction.
Evaluates LLMs' ability to extract and structure information from unstructured text into deep, multi-layer nested JSON formats. It probes format fidelity, field correctness, and structural completeness across varying nesting depths and domains. Use when the user wants to benchmark on DeepJSONEval, or asks about evaluating this task. Reports detailed score.
This evaluation probes the translation quality and contextual consistency of two commercial machine translation systems (DeepL and Supertext) by having professional raters perform blind pairwise comparisons on full documents. It specifically measures whether LLM-based long-context translation yields superior document-level coherence compared to traditional segment-level systems. Use when the user wants to benchmark on Unspecified source documents, or asks about evaluating this task. Reports p...
Evaluates deep learning models for binary malware traffic detection using raw network bytestreams. It probes the model's ability to distinguish benign from malicious network flows or packets without relying on handcrafted domain features. Use when the user wants to benchmark on USTCTFC2016, or asks about evaluating this task. Reports accuracy.
Evaluates mathematical reasoning capabilities on a curated, decontaminated dataset of challenging problems, measuring performance across standardized math competitions and academic benchmarks. Use when the user wants to benchmark on DeepMath-103K, or asks about evaluating this task. Reports accuracy.
Evaluates continuous control reinforcement learning agents on a standardized suite of physics-based simulation tasks. It probes sample efficiency, stability, and performance over long training horizons using uniform action, observation, and reward structures. Use when the user wants to benchmark on DeepMind Control Suite, or asks about evaluating this task. Reports return.
Evaluates deep learning models on a comprehensive suite of protein sequence learning tasks, including function prediction, subcellular localization, protein-protein interaction, epitope/paratope prediction, antibody developability, CRISPR repair outcomes, and protein structure prediction. Use when the user wants to benchmark on Fluorescence, Stability, β-lactamase, Solubility, Subcellular, Binary, PPI Affinity, Yeast, Human PPI, IEDB, PDB-Jespersen, SAbDab-Liberis, TAP, SAbDab-Chen, CRISPR-Le...
Evaluates the throughput, tail latency, and power efficiency of a dynamic scheduling system for at-scale neural recommendation inference across various industry models, hardware platforms, and tail-latency constraints. Use when the user wants to benchmark on Industry-representative recommendation models (DLRM-RMC1/2/3, WND, MT-WND, NCF, DIN, DIEN), or asks about evaluating this task. Reports QPS.
Evaluates an autonomous AI research system's ability to progressively advance state-of-the-art methods across three distinct AI tasks: agent failure attribution, LLM inference acceleration, and AI text detection. It also assesses the scientific quality of the AI-generated research papers through automated and human peer review. Use when the user wants to benchmark on Who&When benchmark, MBPP, AI Text Detection dataset, or asks about evaluating this task. Reports Accuracy, AUROC.
This evaluation probes an end-to-end speech recognition system's ability to accurately transcribe conversational telephone speech and robustly handle background noise without phoneme-level modeling or explicit speaker adaptation. It measures transcription accuracy against ground truth references using standard error rates. Use when the user wants to benchmark on Switchboard Hub5’00 (LDC2002S23), Custom Noisy Speech Test Set, or asks about evaluating this task. Reports word error rate (WER).
Evaluates neural information retrieval models on ad-hoc web search and benchmark datasets by measuring how well they rank relevant documents using segment-level matching and discourse structure modeling. Use when the user wants to benchmark on TREC 2010-2012 Web Track, LETOR 4.0 MQ2008, or asks about evaluating this task. Reports nDCG@20.
Evaluates a deep learning model's ability to reconstruct subglacial bed topography by fusing sparse radar ice thickness measurements with high-resolution surface elevation and ice dynamics data. It probes spatial interpolation accuracy, structural preservation of terrain features, and robustness in data-scarce glacial regions. Use when the user wants to benchmark on Upernavik Isstrøm, or asks about evaluating this task. Reports MAE.
Evaluates trajectory prediction and planning capabilities in high-density urban environments with significant vehicle-to-vulnerable-road-user interactions. It measures prediction accuracy and safety compliance using displacement errors and collision scores. Use when the user wants to benchmark on DeepUrban, or asks about evaluating this task. Reports ADE.
Evaluates the multimodal mathematical reasoning and general multimodal reasoning capabilities of vision-language models. It probes visual perception, step-by-step logical deduction, and cross-domain generalization on K12-level math and broader visual tasks. Use when the user wants to benchmark on Multimodal Math & General Reasoning Benchmarks (WeMath, MathVerse_vision, MathVision, LogicVista, MMMU_VAL, MMMU_Pro_full, M^3CoT), or asks about evaluating this task. Reports accuracy.
Evaluates agentic systems' ability to perform wide-scale information collection and deep multi-hop reasoning simultaneously to fill structured result tables. It probes combinatorial search complexity, tool orchestration, reflection, and context management in real-world information-seeking tasks. Use when the user wants to benchmark on DeepWideBenchmark, or asks about evaluating this task. Reports Success Rate.
This evaluation probes a model's ability to dynamically retrieve and reason over multimodal evidence to verify open-domain claims. It tests zero-shot multimodal fact-checking across text-only, text-image, and out-of-context scenarios, measuring both classification accuracy and the quality of generated justifications. Use when the user wants to benchmark on AVeriTeC, MOCHEG, VERITE, ClaimReview2024+, or asks about evaluating this task. Reports accuracy.
Evaluates industrial defect segmentation models by measuring their ability to accurately localize and classify multiple defect types within complex manufacturing images. It also assesses how well models trained on refined, granular annotations generalize compared to those trained on coarse original annotations, using both pixel-level and image-level quality control metrics. Use when the user wants to benchmark on Defect Spectrum, or asks about evaluating this task. Reports mIoU.
Evaluates a pixel-wise deferral framework for medical image segmentation that dynamically routes uncertain pixels to synthetic or real experts. It measures how well the collaborative system improves segmentation accuracy over strong baselines like MedSAM across diverse organs and imaging modalities. Use when the user wants to benchmark on PROMISE12, LiTS, AMOS22, Chaksu, or asks about evaluating this task. Reports DSC.
Evaluates time series forecasting accuracy and prediction interval calibration using a data-centric cross-similarity approach. It probes the model's ability to aggregate future paths from similar historical reference series to generate point forecasts and uncertainty bounds across different frequencies and historical sample lengths. Use when the user wants to benchmark on M1 and M3 forecasting competitions, or asks about evaluating this task. Reports MASE.
Evaluates the detectability and orbital parameter recovery of long-period exoplanets using simulated Gaia astrometric observations. It quantifies detection significance by comparing the goodness-of-fit between a standard astrometric model and a full orbital model. Use when the user has predictions and gold and needs to compute Δχ².
This benchmark evaluates a model's ability to detect hallucinations in domain-specific question answering systems that use retrieval-augmented generation. It probes whether models can correctly identify when a generated answer contradicts or goes beyond the provided retrieved context, often due to over-reliance on pre-trained knowledge or incomplete retrieval. Use when the user wants to benchmark on DelucionQA, or asks about evaluating this task. Reports Macro F1.
Evaluates the alignment between expected demographic proportions in a target population and the actual representation in a dataset or model outputs. It probes sampling, deployment, and structural biases by quantifying discrepancies across protected attributes. Use when the user has predictions and gold and needs to compute demographic_disparity.
Evaluates a model's ability to comprehend and follow complex, interleaved multimodal instructions that require inferring missing visual details and reasoning across multiple images and text turns. It probes reasoning-aware detail comprehension, image-text alignment, and sensitivity to visual context order. Use when the user wants to benchmark on DEMON, MME, OwlEval, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to learn dense, pose-robust 3D canonical embeddings for human head images. It probes geometric fidelity in point matching, semantic consistency across identities, and robustness to occlusions and extreme poses. Use when the user wants to benchmark on CelebV-HQ, Nersemble, or asks about evaluating this task. Reports MAE.