
Claude Skills by qhjqhj00
github.com/qhjqhj00This evaluation protocol tests the scalability and correctness of SMT-based formal verification for quantized neural networks. It measures how well an SMT model-checking framework can prove safety properties or find counterexamples across different quantization levels, network architectures, and SMT solvers. Use when the user wants to benchmark on Iris dataset, Vocalic dataset, AcasXu benchmark, or asks about evaluating this task. Reports verification_time.
Evaluates the performance of quantum and classical optimization algorithms on intractable combinatorial problems. It measures solution quality, algorithmic success rates, and computational efficiency to track progress toward quantum advantage. Use when the user wants to benchmark on QOBLIB, or asks about evaluating this task. Reports best_objective_value.
Evaluates mathematical reasoning, knowledge-oriented language understanding, and challenging reasoning capabilities of LLMs fine-tuned on domain-specific corpora. Use when the user wants to benchmark on MATH, GSM8K, MMLU, AGIEval, BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.
Measures social bias in medical question-answering systems for pain management by evaluating treatment denial rates across intersectional race-gender profiles. It probes whether AI models exhibit discriminatory prescribing patterns when presented with clinical vignettes containing demographic attributes. Use when the user wants to benchmark on Q-Pain, or asks about evaluating this task. Reports probability_of_no.
Evaluates whether citation growth outpaces publication expansion across leading AI/NLP conferences, measuring the elasticity of scholarly impact relative to scale-driven growth over a decade. Use when the user has predictions and gold and needs to compute QQE.
Evaluates a model's ability to generate semantically equivalent paraphrase sentences from an input question, measuring lexical and semantic overlap with ground truth references. The benchmark probes sentence-level semantic understanding and generative fluency in a question-paraphrase setting. Use when the user wants to benchmark on Quora Question Pairs (QQP), or asks about evaluating this task. Reports BLEU.
Evaluates machine reading comprehension on a low-resource religious domain (Qur'an). It probes a model's ability to extract precise answer spans from Arabic text given a question, testing both exact matching and partial semantic/token overlap. Use when the user wants to benchmark on QRCD, or asks about evaluating this task. Reports pRR.
Evaluates the tradeoff between response fluency and factual attribution in retrieval-augmented conversational LLMs. It measures how well models generate coherent, context-aware responses while correctly grounding answers in provided evidence or dialog history. Use when the user wants to benchmark on QReCC, or asks about evaluating this task. Reports Auto-AIS.
Evaluates a quantum support vector machine (QSVM) with quantum feature selection for binary fraud detection on real-world card payment data. It probes the model's ability to identify fraudulent transactions using a balanced dataset and compares performance against classical feature selection baselines. Use when the user wants to benchmark on Real-world card payment data (Balanced Data Set), or asks about evaluating this task. Reports Accuracy.
Evaluates the calibration of voxel-wise uncertainty estimates in brain tumor segmentation. It measures how effectively uncertainty thresholds filter out incorrect predictions while preserving correct ones, rewarding high confidence in accurate regions and penalizing the loss of correct predictions when filtering uncertain voxels. Use when the user has predictions and gold and needs to compute QU-BraTS unified score.
Evaluates multi-turn, context-dependent question answering where models must track dialog history, resolve coreference, and handle open-ended follow-ups. It specifically probes a system's ability to manage asymmetric knowledge access and correctly identify unanswerable questions within an information-seeking dialogue. Use when the user wants to benchmark on QuAC, or asks about evaluating this task. Reports word-level F1.
Evaluates a drone's ability to autonomously navigate to a target vehicle using marker-free visual servoing. It measures tracking accuracy, localization precision, and flight efficiency in both simulation and real-world environments. Use when the user wants to benchmark on Custom Simulation & Real-World Flight Dataset, or asks about evaluating this task. Reports NormError.
This benchmark evaluates offline reinforcement learning algorithms on real-world quadrupedal locomotion tasks. It probes the policy's ability to accurately track locomotion commands, maintain energy efficiency, and exhibit stability under real-world environmental stochasticity and terrain variations. Use when the user wants to benchmark on Real-World Quadrupedal Locomotion Dataset, or asks about evaluating this task. Reports Return.
This evaluation probes the ability of multi-agent guardrail systems to enforce machine-checkable safety policies over agent trajectories in real-time. It measures how effectively a system detects and blocks unsafe actions while minimizing false positives on enterprise web-agent and malicious behavior benchmarks. Use when the user wants to benchmark on ST-WebAgentBench, AgentHarm, or asks about evaluating this task. Reports Accuracy.
Probes real-time audio-visual reasoning and situated common sense in dialogue. It requires models to resolve deictic references, perform temporal grounding, and integrate evolving visual and auditory streams to answer open-ended questions posed during video playback. Use when the user wants to benchmark on Qualcomm IVD, or asks about evaluating this task. Reports Corr..
Evaluates auditory large language models' ability to perceive and describe low-level speech quality aspects, including noise, distortion, speed, continuity, listening effort, naturalness, and overall quality. It probes both numerical score prediction and natural language reasoning/description generation. Use when the user wants to benchmark on QualiSpeech, or asks about evaluating this task. Reports PCC.
Evaluates a model's ability to classify machine translation outputs as 'good' (zero HTER) or 'bad' (non-zero HTER) for practical post-editing filtering. It probes whether binary classification outperforms thresholded regression for identifying adequate translations in real-world deployment scenarios. Use when the user wants to benchmark on WMT17 QE/QC, or asks about evaluating this task. Reports R@P_t.
This benchmark evaluates a model's ability to comprehend and reason over long documents (2k–8k tokens) to answer multiple-choice questions. It specifically probes whether models can integrate global context rather than relying on local keyword matching or summaries, with a subset (HARD) filtering for questions that require full reading rather than skimming. Use when the user wants to benchmark on QuALITY, or asks about evaluating this task. Reports accuracy.
Compute the QualityWithNoReference metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute QualityWithNoReference, or asks how to score with QualityWithNoReference.
Evaluates open-domain numerical fact-checking by testing how well models can verify claims using decomposed queries and retrieved evidence. It probes the impact of claim decomposition quality on evidence retrieval and downstream NLI-based verification accuracy. Use when the user wants to benchmark on QuanTemp++, or asks about evaluating this task. Reports accuracy.
Evaluates Training Data Attribution (TDA) methods by measuring how accurately they approximate counterfactual training effects, detect mislabeled or shortcut-dependent samples, and support downstream classification tasks. Use when the user wants to benchmark on General TDA Benchmark Suite, or asks about evaluating this task. Reports Linear Datamodeling Score (LDS).
Evaluates a model's ability to fact-check real-world numerical claims containing statistical and temporal expressions. It probes evidence retrieval, claim decomposition, and natural language inference to predict veracity (True, False, or Conflicting). Use when the user wants to benchmark on NumTemp, or asks about evaluating this task. Reports Macro-F1.
This benchmark evaluates a model's ability to classify high-energy physics particle jets as originating from quarks or gluons using pixelized detector data. It probes feature extraction and binary classification performance across different input channel configurations and jet transverse momentum ranges. Use when the user wants to benchmark on Simulated CMS LHC Jet Data (DELPHES), or asks about evaluating this task. Reports AUC.
Evaluates whether rewriting ambiguous queries using answer-free context improves factual QA accuracy compared to standard RAG baselines. It probes a model's ability to leverage disambiguated queries for better retrieval-augmented generation and measures the semantic alignment between rewritten queries and grounding contexts. Use when the user wants to benchmark on HLE-subset, flashrag_fermi, ai_plan, arXiv_2502_17521v1, or asks about evaluating this task. Reports benchmark accuracy.
Evaluates a session-based query suggestion model's ability to rank candidate queries and generate plausible next queries. It probes the model's discriminative ranking capability and its generative quality in capturing user intent and query reformulation patterns. Use when the user has predictions and gold and needs to compute MRR.
Compute Qui-nn/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Qui-nn/accents_unplugged_eval.
Evaluates time series forecasting models on a regime-balanced benchmark stratified by trend, seasonality, and forecastability. It probes a model's ability to handle varying context lengths, forecast horizons, and multivariate dependencies while mitigating domain bias and information leakage. Use when the user wants to benchmark on QuitoBench, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to recommend relevant quotes for a given writing context. It probes cross-lingual recommendation capabilities across English, Standard Chinese, and Classical Chinese by measuring how accurately and highly a model ranks the correct quote among a pool of candidates. Use when the user wants to benchmark on QuoteR, or asks about evaluating this task. Reports MRR.
Evaluates the reliability and accuracy of crowdsourced annotations for Quranic recitation audio. It measures annotator performance against expert labels, assesses inter-rater consistency, and validates an automated label-selection algorithm. Use when the user wants to benchmark on Quranic Audio Dataset, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
This protocol evaluates large language models across core competencies including general knowledge, reasoning, coding, and mathematics. It also assesses multilingual understanding, instruction following, and alignment with human preferences. Use when the user wants to benchmark on MMLU, MMLU-Pro, GPQA, HumanEval, GSM8K, MT-Bench, IFEval, or asks about evaluating this task. Reports accuracy.
This evaluation protocol assesses the speech synthesis and tokenization capabilities of the Qwen3-TTS model. It probes zero-shot voice cloning, multilingual and cross-lingual generation, controllable voice design, and long-form audio stability across multiple languages. Use when the user wants to benchmark on CommonVoice & Fleurs, LibriSpeech test-clean, Seed-TTS test set, TTS multilingual test set, CV3-Eval, InstructTTSEval, or asks about evaluating this task. Reports WER.
Evaluates complex reasoning capabilities of LLMs and MLLMs on graduate-level, multi-disciplinary academic questions in both English and Chinese. It probes the models' ability to handle rigorous, curriculum-based problems requiring extended chain-of-thought reasoning. Use when the user wants to benchmark on R-Bench-T, R-Bench-M, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates the accuracy of a regression framework (R-ICE) in estimating prompt-level inference carbon and energy emissions for LLMs using only token counts and publicly available performance data. It probes whether runtime can be reliably modeled as a piecewise linear function of input and output tokens without intrusive monitoring or architecture details. Use when the user wants to benchmark on HELM, or asks about evaluating this task. Reports average prediction error.
Evaluates LLMs' ability to judge safety risks in multi-turn agent interactions by classifying whether a given interaction record poses a safety risk. It probes risk perception and binary safety classification under zero-shot and few-shot prompting conditions, with and without explicit risk descriptions. Use when the user wants to benchmark on R-Judge, or asks about evaluating this task. Reports F1.
Evaluates the semantic alignment and retrieval capability of scene graph generation models and image-scene graph similarity frameworks. It measures how well generated or ground-truth scene graphs can retrieve corresponding images compared to other images in a dataset. Use when the user wants to benchmark on Visual Genome, Open Images V6, or asks about evaluating this task. Reports R-Precision@K.
Evaluates a model's ability to perform unsupervised anomaly detection on multi-agent traffic trajectories in urban environments. It probes frame-wise recognition of rare and abnormal driving behaviors, including both individual maneuvers and context-dependent interactions between agents and static map features. Use when the user wants to benchmark on R-U-MAAD, or asks about evaluating this task. Reports frame-wise anomaly detection accuracy.
Probes LLMs' ability to autonomously generate and execute code for complex reasoning and planning tasks across logic, spatial, order, optimization, search, and math domains. It evaluates how well models can iteratively explore, optimize, and self-check solutions in multi-turn code execution environments. Use when the user wants to benchmark on SymBench, Big-Bench-Hard, Reasoning-Gym, or asks about evaluating this task. Reports exact match or constraint check.
Compute the r2_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute r2_score, or asks how to score with r2_score.
Evaluates retrieval models on reasoning-driven medical tasks where document relevance is determined by alignment with inferred clinical diagnoses or multi-step reasoning paths rather than lexical or semantic overlap. Covers three task types—Q&A reference, clinical evidence, and clinical case retrieval—spanning eight medical sub-domains. Use when the user wants to benchmark on R2MED, or asks about evaluating this task. Reports nDCG@10.
Evaluates the ability to detect incorrect answers in chain-of-thought reasoning by analyzing inconsistencies across multiple reasoning paths. It measures how well intermediate steps can predict the correctness of a final answer without relying solely on the answer itself. Use when the user wants to benchmark on R2PE, or asks about evaluating this task. Reports Discernibility Score (DS).
Compute the R2Score metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute R2Score, or asks how to score with R2Score.
This evaluation probes the in-context learning (ICL) capabilities of reinforcement learning agents across diverse environments, including grid-worlds, robotics simulators, and video games. It measures how effectively an agent can leverage retrieved past experiences to improve its policy over consecutive interaction trials without weight updates. Use when the user wants to benchmark on Dark-Room, Dark Key-Door, MazeRunner, Meta-World, DMControl, Procgen, or asks about evaluating this task. Rep...
Evaluates high-speed autonomous driving capabilities including precise localization, long-range object detection and tracking, and robust mapping/SLAM under extreme dynamic conditions (up to 170 mph). The protocol benchmarks how well models maintain accuracy and latency when processing multi-modal sensor data at racing speeds where motion blur, sensor dropout, and rapid ego-motion are prevalent. Use when the user wants to benchmark on RACECAR, or asks about evaluating this task. Reports Avera...
Probes multi-sport vision capabilities by evaluating unified ball tracking, racket pose estimation, and dynamic trajectory prediction across table tennis, tennis, and badminton. It tests both static perception and temporal modeling of human-object interactions. Use when the user wants to benchmark on RacketVision, or asks about evaluating this task. Reports ball tracking.
Evaluates the robustness of image anomaly detection models against real-world imaging distortions, including free viewpoints, uneven illumination, and motion blur. It measures how well unsupervised and zero-shot methods localize and classify anomalies on industrial work platforms with foreign objects. Use when the user wants to benchmark on RAD, or asks about evaluating this task. Reports AUROC.
Evaluates multi-modal large language models on weather radar forecast quality analysis, specifically testing their ability to perform quantitative rating of radar frames/sequences and generate qualitative assessment reports. It probes domain-specific meteorological understanding, temporal pattern evolution tracking, and alignment with expert judgment. Use when the user wants to benchmark on RQA-70K, or asks about evaluating this task. Reports accuracy.
Evaluates a vision-language model's ability to perform joint radiological diagnosis, abnormality detection, and multi-target segmentation on X-ray and CT images. It probes the model's capacity for open-ended visual question answering, precise pixel-level mask generation, and robustness to label-imbalanced medical data. Use when the user wants to benchmark on RadDiagSeg-D, VQA-RAD, SLAKE, or asks about evaluating this task. Reports F1, Dice.
Evaluates an AI pipeline's ability to generate clinically accurate radiology reports from 3D CT scans by detecting tumors, measuring their size, localizing them within organ sub-segments, and staging cancers. It also assesses the textual similarity and diagnostic utility of generated reports compared to ground-truth clinical notes. Use when the user wants to benchmark on AbdomenAtlas 3.0, or asks about evaluating this task. Reports Tumor Detection Sensitivity & Specificity.
Evaluates the training efficiency and novel-view synthesis or reconstruction quality of neural radiance field methods by reducing the number of rays sampled during volume rendering. It measures how well adaptive ray allocation preserves rendering accuracy while accelerating convergence across diverse 3D scene benchmarks. Use when the user wants to benchmark on Realistic Synthetic 360°, Light Field (LF), LLFF, Tanks and Temples (T&T), Real-World 360°, DTU, or asks about evaluating this task. R...
Evaluates the accuracy of a numerical vector radiative transfer code by comparing its outputs for scattering phase functions, polarization, and albedo against established reference results from prior literature. Use when the user has predictions and gold and needs to compute disk-integrated total degree of polarization.