
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multilingual and multidomain aspect-based sentiment analysis by predicting continuous valence-arousal (VA) scores alongside aspect, opinion, and category extraction. It probes a model's ability to perform fine-grained dimensional sentiment regression and structured information extraction across diverse languages and domains. Use when the user wants to benchmark on DimABSA, or asks about evaluating this task. Reports RMSE_VA, cF1.
Evaluates a model's ability to generate syntactically and semantically correct SQL queries from natural language questions across diverse database schemas. It probes schema linking, handling of complex query structures (joins, nested subqueries, aggregations), and the capacity for iterative self-correction when initial generations fail. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
This benchmark evaluates the robustness of network intrusion detection models against distribution shifts between different network environments. It specifically probes cross-domain generalization by training on one NetFlow dataset and testing on another, measuring how well domain-invariant feature extraction mitigates performance degradation when facing unseen attack distributions. Use when the user wants to benchmark on NFv2-UNSW-NB15, NFv2-CIC-2018, or asks about evaluating this task. Repo...
This benchmark evaluates Vision-Language Models on fine-grained visual discrimination, nutritional quantification from images, and complex food-related visual question answering. It probes the models' ability to fuse multi-view imagery, perform volumetric reasoning, and avoid parametric knowledge biases when identifying dishes and estimating macronutrients. Use when the user wants to benchmark on DiningBench, or asks about evaluating this task. Reports Accuracy, MAPE.
Evaluates the cross-task generalizability of the DINOv2 vision foundation model on medical image analysis tasks, specifically disease classification and organ segmentation across X-ray, CT, and MRI modalities. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, SARS-CoV-2, Brain Tumor, Montgomery County (MC), AMOS, MSD Heart, MSD Hipp, MSD Spleen, or asks about evaluating this task. Reports AUROC.
Evaluates the quality of frozen self-supervised visual features by training a simple linear classifier on top of them across diverse image and video understanding tasks, probing generalization, robustness, and instance-level recognition capabilities. Use when the user wants to benchmark on ImageNet-1k, ImageNet-V2, ImageNet-ReaL, iNaturalist, Places205, UCF-101, Kinetics-400, Something-Something v2, Oxford/Paris, ImageNet-A, ImageNet-R, ImageNet-C, Sketch, or asks about evaluating this task. ...
Evaluates the cross-domain generalization and scaling behavior of a natural-image pre-trained vision transformer (DINOv3) across diverse medical imaging modalities, including 2D/3D classification and segmentation tasks. Use when the user wants to benchmark on NIH-14, RSNA-Pneumonia, Camelyon16, Camelyon17, BCNB, Kvasir-Capsule, AutoLaparo, EndoVis18, EDD 2020, CT-RATE, Medical Segmentation Decathlon (MSD), CREMI, AC3/4, AutoPET-II, HECKTOR 2022, or asks about evaluating this task. Reports AUC...
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
Evaluates machine translation systems on four discourse phenomena: anaphora resolution, lexical consistency, coherence/readability, and discourse connectives. It probes whether context-aware models can maintain discourse-level quality and consistency across different language pairs beyond standard n-gram overlap. Use when the user wants to benchmark on DiP Benchmark, or asks about evaluating this task. Reports BLEU.
Evaluates distant speech recognition and noise robustness by measuring word error rate on close-talk and far-field recordings of natural dinner conversations. It tests the model's ability to handle uncontrolled acoustic conditions, background music, and spatially diverse microphone placements. Use when the user wants to benchmark on DiPCo, or asks about evaluating this task. Reports WER.
Assesses machine translation quality by measuring the proportion of source words that have strong lexical co-occurrence evidence in the training corpus. It evaluates whether data-driven lexical transfer fidelity correlates with standard translation quality metrics like BLEU. Use when the user has predictions and gold and needs to compute DE Score.
Evaluates whether the model refuses or safely handles requests for disallowed content (e.g., hate speech, illicit advice, personal data, self-harm, sexual/exploitative material) across standard and production-like multiturn conversations. Use when the user wants to benchmark on Production Benchmarks, or asks about evaluating this task. Reports not_unsafe.
Evaluates large vision-language models' ability to assess disaster damage from remote sensing imagery. It probes capabilities in multi-sensor (optical/SAR) understanding, object counting, relational reasoning, and generating professional disaster response reports. Use when the user wants to benchmark on DisasterM3, or asks about evaluating this task. Reports accuracy (%).
Evaluates information retrieval models on disaster management queries across multiple intents and categories. It measures how well models retrieve relevant passages from a large-scale domain-specific corpus under both exact and approximate nearest neighbor search settings. Use when the user wants to benchmark on DisastIR, or asks about evaluating this task. Reports NDCG@10.
This benchmark evaluates image copy detection systems under realistic, adversarial conditions. It probes a model's ability to match transformed query images against a large reference database while resisting geometric, color, overlay, and deepfake manipulations. The setup emphasizes scalability and robustness in a high-false-positive-rate, needle-in-haystack search regime. Use when the user wants to benchmark on DISC21, or asks about evaluating this task. Reports micro Average Precision.
Evaluates the ability of fine-tuned LLMs to generate clinically accurate, complete, and readable discharge summaries for cardiac patients from raw medical records. It probes domain-specific medical summarization, factual consistency, and adherence to clinical documentation standards. Use when the user wants to benchmark on Cardiology Clinical Dataset, or asks about evaluating this task. Reports Accuracy.
This protocol evaluates how well a metamodel can predict the benchmark performance of unseen models using a highly condensed subset of test samples. It probes the trade-off between evaluation cost reduction and the fidelity of accuracy estimation and model ranking preservation across language and vision benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, Winogrande, ARC, ImageNet-1k, or asks about evaluating this task. Reports MAE, Spearman rank correlation.
Evaluates whether sentence representations capture discourse-aware semantics by testing performance on tasks involving sentence ordering, discourse relations, and coherence across multiple domains. Use when the user wants to benchmark on DiscoEval, or asks about evaluating this task. Reports accuracy.
Evaluates speech enhancement models in extremely low SNR conditions by measuring noise suppression, speech quality preservation, and intelligibility using both objective metrics and subjective listening tests. Use when the user wants to benchmark on Low-SNR Dataset, VB-DMD Dataset, DNS Non-Reverb Test Dataset, DNS Real Recordings, or asks about evaluating this task. Reports PESQ.
This benchmark evaluates the ability of discrete speech models to unsupervisedly discover phoneme inventories and capture phonemic contrasts. It measures how well predicted units align with gold phonemes through phonetic similarity, recognition error, and temporal segmentation accuracy across multiple languages. Use when the user wants to benchmark on discoPhon, or asks about evaluating this task. Reports PNMI.
Evaluates whether text-to-speech systems can correctly realize discourse-dependent word-level stress based on contrasting contexts. It probes the model's ability to adapt prosodic emphasis dynamically rather than relying on fixed sentence-internal stress patterns. Use when the user wants to benchmark on CAST, or asks about evaluating this task. Reports Pair-Correct.
Evaluates the ability of machine learning classifiers to distinguish rare signal events from dominant Standard Model backgrounds in simulated high-energy physics collisions, and their capacity to accurately estimate signal fractions and discovery sensitivity via unbinned template fits. Use when the user has predictions and gold and needs to compute discovery-sensitivity.
Evaluates machine translation systems on discourse-level coherence and terminological precision in expert domains. It probes the model's ability to maintain long-form text consistency and handle domain-specific language beyond sentence-level translation. Use when the user wants to benchmark on DiscoX, or asks about evaluating this task. Reports Metric-S.
This benchmark evaluates discrete audio tokenizers across speech, music, and general audio domains. It probes their ability to preserve acoustic fidelity during compression and decompression, as well as their effectiveness when used as inputs for downstream discriminative and generative audio tasks. Use when the user wants to benchmark on LibriSpeech test-clean, MUSDB, Audioset test-set, DASB Benchmark, or asks about evaluating this task. Reports SDR.
Evaluates a model's ability to predict user check-in behavior (CTR) in location-based recommendation by disentangling sequential and geographical influences. It probes how well the model handles data sparsity and cold-start scenarios using real-world POI interaction logs. Use when the user wants to benchmark on Foursquare Tokyo, Foursquare New York, Meituan, or asks about evaluating this task. Reports AUC.
Evaluates pre-trained language models' ability to reason about diseases by mapping symptoms, treatments, tests, procedures, and terminology to disease names. It isolates medical reasoning types and uses adversarial negative examples to prevent knowledge leakage. Use when the user wants to benchmark on DisKnE, or asks about evaluating this task. Reports F1 score.
Probes a model's ability to perform dynamic aspect-based summarization on disordered, non-sequential texts where sentences from multiple sources are shuffled. It tests whether the model can cluster fragmented content by underlying topics or aspects and generate precise, coherent summaries without relying on original sentence order. Use when the user wants to benchmark on D-CnnDM, D-WikiHow, or asks about evaluating this task. Reports Human Evaluation (Coherence, Consistency, Fluency, Relevanc...
Evaluates a distributed multi-agent routing system's ability to autonomously assign queries to the most appropriate LLM based on intrinsic self-assessment. It probes the trade-off between routing accuracy and inference cost across diverse mathematical, commonsense, and reading comprehension benchmarks. Use when the user wants to benchmark on GSM8K, ARC, MMLU, RACE_HIGH, OpenbookQA, DROP, CosmosQA, SQuAD, HellaSwag, HeadQA, or asks about evaluating this task. Reports utility.
Evaluates vision-language models' ability to perform fine-grained low-level visual perception by identifying both the specific type of image distortion and its severity level from a single image. It probes whether models rely on direct perceptual pattern matching or struggle with subtle severity discrimination. Use when the user wants to benchmark on DistortBench, or asks about evaluating this task. Reports Acc..
Evaluates link recommendation models on social networks by measuring how well they predict future connections while respecting individual users' diversity preferences across profile dimensions. Use when the user wants to benchmark on Large-scale social network datasets, or asks about evaluating this task. Reports F1 Score.
Evaluates the execution speed and hardware resource utilization of three open-source deep learning frameworks (TensorFlow, Theano, CNTK) across standard computer vision, NLP, and custom datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, IMDB, Self-Driving Car, Penn TreeBank, or asks about evaluating this task. Reports processing_time.
Evaluates the robustness of deep learning models trained on different frameworks (TensorFlow, Theano, Torch) against adversarial attacks. It measures how well models maintain correct predictions when subjected to white-box, black-box, and decision-based perturbations. Use when the user wants to benchmark on MNIST, CIFAR-10, or asks about evaluating this task. Reports robustness indicator R(m_i).
Evaluates the predictive accuracy and computational efficiency of deep learning models for urban traffic forecasting. It benchmarks grid-based, graph-based, and multivariate time-series architectures on standard traffic datasets to compare their ability to capture spatiotemporal dependencies. Use when the user wants to benchmark on BikeNYC-I, TaxiNYC, TaxiBJ, METR-LA, PeMS-BAY, PEMSD7M, or asks about evaluating this task. Reports MAE.
This protocol evaluates information retrieval systems by measuring their ranking effectiveness on passage retrieval tasks using both original seed queries and LLM-generated query variants aligned with specific demographic or textual profiles. It probes whether retrieval systems perform consistently across diverse user personas and query transformations, revealing potential disparities in system behavior and ranking stability. Use when the user wants to benchmark on DL21 & DL22 (TREC Deep Lear...
Evaluates the accuracy of a composable benchmark generation framework in estimating deep learning model inference latency on CPUs, and measures the computational speedup achieved by benchmarking unique layer sequences instead of running full end-to-end models. Use when the user wants to benchmark on 50 DL Models, or asks about evaluating this task. Reports Normalized Latency.
Evaluates the ability of multiple-instance learning models to classify six key pathological indicators (cholestasis, portal fibrosis, inflammation, steatosis, macrovesicular steatosis, hepatocellular ballooning) from donor liver whole slide histopathological images. Use when the user wants to benchmark on DLiPath, or asks about evaluating this task. Reports AUC, Accuracy, Precision, Recall, F1-Score.
Evaluates the effectiveness and efficiency of Diffusion LLMs versus Autoregressive LLMs across the software development lifecycle. It probes code generation accuracy, binary defect detection, automated program repair, and cross-file issue resolution, while measuring generation throughput and latency. Use when the user wants to benchmark on HumanEval, Mercury, Devign, Bears, Defects4J, SWE-bench, or asks about evaluating this task. Reports Pass@K.
Evaluates ranking models' ability to align machine-generated relevance with fine-grained user intents, particularly for ambiguous or multi-intent queries. It also measures the diversity of search results when multiple user intents are fused into a single ranking. Use when the user wants to benchmark on DL-MIA, or asks about evaluating this task. Reports α-nDCG@10.
Evaluates document language understanding across four core capabilities: classification, structural analysis, information extraction, and transcription. It probes models' ability to handle long documents, complex hierarchical structures, and dispersed knowledge spread across large contexts. Use when the user wants to benchmark on Hyperpartisan, ContractNLI, ECOM, RR, GUM, LitBank, NarrativeQA, SummScreen, GovReport, Qasper, or asks about evaluating this task. Reports F1.
Evaluates long-horizon physical state prediction accuracy of a neural motion simulator in continuous control environments, and measures its effectiveness for zero-shot reinforcement learning by comparing prediction horizons and minimal training step requirements. Use when the user wants to benchmark on DM Control, or asks about evaluating this task. Reports MSE loss.
Evaluates sample efficiency, asymptotic performance, and generalization of reinforcement learning agents on high-dimensional continuous control tasks with varying observation modalities (state, pixels, multi-modal) and reward structures (dense, sparse, goal-conditioned). Use when the user wants to benchmark on DeepMind Control Suite (DMControl), Meta-World v2, or asks about evaluating this task. Reports Cumulative Episode Return.
Evaluates a vision-language model's ability to generate clinically accurate and linguistically fluent mammography reports from multi-view breast images. It probes both natural language generation quality and domain-specific diagnostic reasoning, specifically BI-RADS categorization and breast density assessment. Use when the user wants to benchmark on DMID, or asks about evaluating this task. Reports BI-RADS Accuracy.
Evaluates genomic foundation models on multiple biological prediction tasks, including regulatory element detection, splicing, and variant-disease association, to measure their ability to capture functional DNA sequences and SNP effects. Use when the user wants to benchmark on Promoter detection, Core promoter detection, TF binding detection, Splicing detection, lenti-MPRA K562, SNP-to-disease association, or asks about evaluating this task. Reports F1 score.
Probes the inference latency and execution time variability of four industrial DNN models across heterogeneous cloud compute instances (CPU, GPU, and inference-optimized) on AWS and Chameleon Cloud. It evaluates how hardware heterogeneity impacts real-time processing requirements for safety-critical industrial applications. Use when the user wants to benchmark on Fire Detection Dataset, Oil Spill Detection Dataset, UCI Human Activity Recognition Using Smartphones, Marmousi 2 Dataset, or asks ...
Evaluates the scalability and correctness of DNN verification tools by measuring their ability to prove safety or robustness properties within a strict time limit across diverse network architectures and property types. Use when the user wants to benchmark on VNN-COMP'22 & MNIST_GDVB, or asks about evaluating this task. Reports verification_success_rate.
Evaluates the perceptual quality and intelligibility of deep noise suppression models under real-world, non-stationary noise conditions. It specifically probes whether models generalize from synthetic training data to real-world acoustic environments. Use when the user wants to benchmark on DNS Challenge Dataset, or asks about evaluating this task. Reports ITU-T P.808.
DO-Bench probes object hallucination in vision-language models by disentangling it into two distinct failure mechanisms: prior-dominated (textual priors overriding visual evidence) and perception-limited (weak visual grounding causing false denials). It uses controlled within-image interventions to measure how models respond to strengthened contextual priors and concentrated visual evidence, revealing heterogeneous failure modes that aggregate accuracy masks. Use when the user wants to benchm...
This benchmark evaluates the safety and harmlessness of language model responses by measuring the proportion of outputs that avoid generating harmful content across various risk categories. It specifically probes a model's ability to refuse or safely handle prompts designed to elicit dangerous, illegal, or unethical outputs. Use when the user wants to benchmark on Do_Not_Answer, or asks about evaluating this task. Reports proportion of harmless responses.
This benchmark isolates and evaluates core visual perception capabilities in multimodal large language models (MLLMs) through programmatically generated tasks inspired by human psychology. It probes figure-ground discrimination, spatial relations, visual form constancy, and perceptual brittleness across 2D and 3D settings with controlled difficulty levels. Use when the user wants to benchmark on Do You See Me, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of deep convolutional neural networks and hand-crafted feature methods to classify scanned document images into predefined categories and retrieve semantically similar documents from a large corpus. Use when the user wants to benchmark on SmallTobacco, BigTobacco, or asks about evaluating this task. Reports classification accuracy.