
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the accuracy, latency, and energy efficiency of a homomorphic inference accelerator performing SVM classification on encrypted data under intermittent power constraints. Use when the user wants to benchmark on MNIST, Human Activity Recognition, ADULT, or asks about evaluating this task. Reports accuracy.
Evaluates object detection models on chest X-rays for identifying critical retained foreign objects (RFOs) like sponges and needles. It probes both classification accuracy and precise localization of rare medical anomalies under data-scarce conditions. Use when the user wants to benchmark on Hopkins RFOs Bench, or asks about evaluating this task. Reports ACC.
Evaluates few-shot node classification on text-attributed graphs using self-supervised preference tuning. It probes the model's ability to leverage graph topology and anchor labels at inference without any supervised training, measuring both classification accuracy and inference efficiency. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy (%).
Evaluates user behavior modeling capabilities across temporal generalization, cross-domain prediction, and unseen-user scenarios. It probes how well recommendation models and LLMs can generalize to out-of-distribution users and future time periods using real-world interaction sequences. Use when the user wants to benchmark on Amazon Reviews (HORIZON Benchmark), or asks about evaluating this task. Reports NDCG@K.
Probes long-horizon personalization and belief-update capability. It tests whether models can track evolving user preferences across ~6 months of conversation history and correctly select responses aligned with updated preferences, rather than anchoring on outdated values. Use when the user wants to benchmark on HorizonBench, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to perform multi-hop question answering by reasoning across multiple documents. It specifically probes explainability through supporting fact prediction and tests robustness against distractor paragraphs and large-scale retrieval contexts. Use when the user wants to benchmark on HotpotQA, or asks about evaluating this task. Reports F1.
Evaluates an agent's ability to detect misplacements of objects in indoor scenes and plan optimal rearrangement placements for carryable objects based on scene context and affordances. It probes commonsense reasoning about object-receptacle relationships and ranking quality under varying contextual cues. Use when the user wants to benchmark on Tidybot benchmark, Context-oriented benchmark (HSSD 200), or asks about evaluating this task. Reports NDCG@8.
Evaluates retrieval and reasoning over housing statutes, requiring models to connect queries to lexically distant legal texts and answer standardized Yes/No or categorical questions. Use when the user wants to benchmark on Housing Statute QA, or asks about evaluating this task. Reports Recall@10.
Probes a model's ability to perform many-hop fact verification by retrieving supporting evidence from multiple Wikipedia articles and determining whether a given claim is supported or not supported. It specifically tests long-range dependency reasoning, coreference resolution, and the ability to avoid semantic matching shortcuts that degrade as hop count increases. Use when the user wants to benchmark on HOVER, or asks about evaluating this task. Reports claim verification accuracy.
Compute hpi-dhc/FairEval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hpi-dhc/FairEval.
Evaluates the quality of the HPLT v2 multilingual corpus by training downstream models (masked language models, generative LMs, and MT systems) and measuring their performance on standard linguistic, natural language understanding, and machine translation benchmarks. Use when the user wants to benchmark on Universal Dependencies (UD) treebanks, WikiAnn, FLORES-200, or asks about evaluating this task. Reports BLEU.
Evaluates hallucinations in large vision-language models by probing their ability to correctly identify object existence and ground fine-grained attributes (color, material, shape) to specific objects in an image. Use when the user wants to benchmark on H-POPE, or asks about evaluating this task. Reports Accuracy.
Evaluates a vision-based framework for classifying object surfaces and recognizing human actions, and tests its integration into a human-robot collaboration controller for ergonomic task execution and subjective workload assessment. Use when the user wants to benchmark on HRI30, Custom Surface Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates multilingual mathematical reasoning capability, specifically probing whether models can comprehend and solve Korean math problems by leveraging English-as-pivot reasoning to bridge cross-lingual comprehension gaps. Use when the user wants to benchmark on HRM8K, or asks about evaluating this task. Reports pass@1.
Evaluates a wavelet-inspired multi-resolution transformer across five linguistic granularities, from character morphology to discourse reasoning. It probes the model's capacity for hierarchical composition, long-range dependency modeling up to 16K tokens, and computational efficiency compared to standard transformers. Use when the user wants to benchmark on WikiMorpho, IMDB-BYTE, WordNet Hypernymy (WN-Hyper), SentEval Word Similarity Suite, GLUE Benchmark, SuperGLUE, Long Range Arena (LRA), W...
Evaluates audio spoof detection models on a newly constructed hybrid spoofing benchmark. It probes robustness against complex, real-world adversarial conditions including mixed-source speech, environmental noise, channel filtering, and compression artifacts. Use when the user wants to benchmark on Hybrid Spoofed Audio Dataset (HSAD), or asks about evaluating this task. Reports Accuracy.
This evaluation probes the ability of audio classification models to detect and distinguish between genuine human speech, AI-cloned speech, AI-generated speech, and complex hybrid compositions that mix human and synthetic segments. It specifically tests robustness against multi-source spoofing attacks and real-world signal degradations like environmental noise, channel filtering, and codec compression. Use when the user wants to benchmark on ASVspoof 2019 Logical Access (LA), Proposed Hybrid ...
This benchmark evaluates deep search agents' ability to perform multi-hop reasoning across hierarchical tariff rules to predict 10-digit Harmonized System Codes (HSCode) from noisy product descriptions and images. It probes rule-based reasoning, agentic knowledge utilization, and handling of vague or implicit classification logic. Use when the user wants to benchmark on HSCodeComp, or asks about evaluating this task. Reports 10-digit accuracy.
Evaluates foundational models' ability to detect social errors and competencies, identify specific social attributes, reason about sequential interaction flow (pre/post conditions), and generate rationales and corrective actions in human-robot interaction scenarios. Use when the user wants to benchmark on HSRI, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to detect HTTP-based Trojan malware in network traffic by analyzing hierarchical spatio-temporal features. It probes the model's binary classification accuracy, robustness to training set class imbalance, and cross-dataset generalization performance. Use when the user wants to benchmark on BTHT-2018, ISCX-2012, or asks about evaluating this task. Reports F1.
Compute huanghuayu/multiclass_brier_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of huanghuayu/multiclass_brier_score.
Evaluates Chinese medical question-answering capabilities through retrieval and generation tasks. It probes domain-specific knowledge retrieval from large pools and tests generative models on producing accurate, long-form medical answers. Use when the user wants to benchmark on Huatuo-26M, or asks about evaluating this task. Reports Recall@5.
This benchmark evaluates anti-spoofing systems for synthetic speech detection, with a specific focus on prosody-awareness, emotional/expressive spoofing, and cross-lingual robustness. It tests whether models can distinguish real speech from TTS, VC, and adversarial attacks across diverse channel conditions and languages. Use when the user wants to benchmark on ASVspoof 2019 (LA), ASVspoof 2021 (LA), ASVspoof 2024 (Track 1), EmoFake, Mixed Emotions, ADD 2022 (Track 1), HABLA, or asks about eva...
Evaluates the computational efficiency and energy cost of NLP models across pretraining, fine-tuning, and inference phases. It measures the time and monetary cost required to reach predefined performance thresholds on standard NLP tasks, normalized against a BERT-Large baseline. Use when the user wants to benchmark on CoNLL 2003, MNLI, SST-2, or asks about evaluating this task. Reports efficiency score.
Evaluates a unified vision-language model's capability to perform holistic medical understanding across text-only queries, 2D/3D medical images, and surgical videos. It probes visual question answering, radiology report generation, clinical reasoning, and multilingual medical dialogue. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, IU X-ray, MedMNIST-2D, M3D, 3D-RAD, AMOS-MM, MedFrameQA, Cholec80-VQA, EndoVis18-VQA, PSI-AVA-VQA, SurgeryVideoQA, MMedBench, RareBench, HealthBench,...
Evaluates a video foundation model's ability to recognize and classify human actions across diverse, real-world, and benchmark video datasets. It tests generalization from self-supervised pre-training on unstructured social media content to structured action recognition tasks. Use when the user wants to benchmark on Kinetics-400, Something-Something V2, UCF-101, HMDB51, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal models' ability to understand and classify diverse psychological and social behaviors (e.g., emotion, sarcasm, depression, intent) across text, audio, and video inputs. It also tests transfer learning capabilities to held-out datasets and the impact of adding behavioral descriptors. Use when the user wants to benchmark on Human Behavior Atlas, or asks about evaluating this task. Reports Unified behavioral metrics.
Evaluates a multi-agent system's ability to generate coherent, context-aware, and emotionally expressive dialogue scripts in simulated human communication scenarios with varying numbers of roles. Use when the user wants to benchmark on Human-Communication Simulation Benchmark, or asks about evaluating this task. Reports Consistency Score.
This evaluation probes the accuracy-cost tradeoff of AI coding agents by measuring how often generated solutions pass test cases relative to the actual inference cost required. It highlights whether complex agent architectures provide genuine performance gains over simple retry baselines when compute expenses are accounted for. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports accuracy.
Evaluates a code generation model's ability to produce correct, executable Python functions from docstrings and function signatures. It measures whether the generated code passes all provided unit tests for each programming problem. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports functional accuracy.
Evaluates text generation models across multiple NLP tasks using standardized human annotation, focusing on reproducibility, annotator quality detection, and scalar scoring of qualities like fluency and correctness. Use when the user has predictions and gold and needs to compute human scores.
Evaluates abstractive summarization quality by measuring human preference over reference summaries and rating outputs across coverage, accuracy, coherence, and overall quality. It also benchmarks how well learned reward models and automatic metrics correlate with human judgments. Use when the user wants to benchmark on Reddit TL;DR, CNN/DailyMail, or asks about evaluating this task. Reports preference score.
Evaluates the accuracy and inference speed of optical flow estimation networks specifically for human motion. It probes a model's ability to capture fine-grained, small-scale displacements typical of human limbs and body parts against complex or layered backgrounds. Use when the user wants to benchmark on Human Flow, or asks about evaluating this task. Reports AEPE.
Probes whether text-to-speech systems can perceptually deceive human listeners into believing synthetic speech is real. It measures the gap between traditional preference scores (CMOS/MUSHRA) and actual indistinguishability, highlighting how prompt expressivity and model type affect deception capability. Use when the user has predictions and gold and needs to compute Human Fooling Rate (HFR).
Evaluates the ability of generative models to synthesize realistic and diverse human motion sequences from text descriptions, as well as their capacity to complete short-term motion sequences given a partial initialization. Use when the user wants to benchmark on Human3.6M (H3.6M), CMU Mocap, or asks about evaluating this task. Reports Inception Score.
This evaluation protocol assesses the accuracy of human and hand pose estimation models in localizing anatomical keypoints on images. It probes the model's ability to handle varying instance scales, occlusion levels, and joint visibility by measuring localization error against ground truth annotations. Use when the user wants to benchmark on MS COCO, MPII Human Pose, RHD, or asks about evaluating this task. Reports AP@OKS.
Evaluates a vision-language model's ability to understand and generate detailed descriptions of human-centric scenes, answer open- and closed-set questions about them, recognize facial attributes, and ground textual references to human objects in images. Use when the user wants to benchmark on HumanCaptionHQ, HumanVQA, FaceC, CelebA, LFWA, RefCOCO, or asks about evaluating this task. Reports semantic similarity.
Evaluates the ability of models to forecast future 3D human poses over a 1-second horizon given a short 40ms observation window. It probes long-term temporal prediction and personalization to individual-specific motion patterns. Use when the user wants to benchmark on Human3.6M, or asks about evaluating this task. Reports MPJE.
Evaluates the generalization and task-agnostic representation learning of human-centric vision models across six diverse downstream tasks. It probes how well a model trained on a large, multi-task human-centric corpus can adapt to in-distribution, out-of-distribution, and completely unseen human perception tasks. Use when the user wants to benchmark on HumanBench, or asks about evaluating this task. Reports mAP/mIoU/mA/Top1/pACC/MR/MSE/EPE.
Evaluates instruction-based image editing models on their ability to modify source images according to textual prompts, with and without provided segmentation masks. It measures pixel-level fidelity, image quality, and text-image alignment across diverse editing categories such as add, remove, replace, action, counting, and relation. Use when the user wants to benchmark on HumanEdit, or asks about evaluating this task. Reports CLIP-T.
Evaluates a model's ability to generate correct, executable Python code from natural language function descriptions and signatures. It measures functional correctness by checking if generated code passes hidden unit tests. Use when the user wants to benchmark on HumanEval, or asks about evaluating this task. Reports pass@1.
Evaluates a model's ability to generate correct Python code for programming tasks and iteratively refine it using execution feedback or simulated human guidance. It measures both initial code generation quality and the effectiveness of a multi-turn debugging loop under strict runtime and edge-case constraints. Use when the user wants to benchmark on HumanEval, MBPP, HumanEval+, MBPP+, or asks about evaluating this task. Reports pass@1.
This benchmark evaluates large multimodal models' ability to perform high-level visual reasoning over complex diagrams in coding contexts. It specifically probes spatial transformations, topological relationships, and dynamic pattern understanding by requiring models to translate visual information into executable code. Use when the user wants to benchmark on HumanEval-V, or asks about evaluating this task. Reports pass@1.
Evaluates multilingual code generation and code translation capabilities across five programming languages (C++, Java, JavaScript, Go, Python). It measures functional correctness by executing generated code against a suite of test cases for each problem. Use when the user wants to benchmark on HumanEval-X, or asks about evaluating this task. Reports pass@k.
Evaluates a model's ability to debug and fix buggy code by generating corrected implementations that pass provided unit tests. It probes code repair capabilities across multiple programming languages. Use when the user wants to benchmark on HumanEvalFix, or asks about evaluating this task. Reports pass rate.
Evaluates imitation learning and vision-language-action policies on open-world humanoid manipulation. It probes robustness to high-dimensional action spaces, multimodal sensor fusion, and fine-grained visuospatial perception across locomotion, tool use, and precise manipulation tasks. Use when the user wants to benchmark on Humanoid Everyday, or asks about evaluating this task. Reports success rate.
Evaluates a unified state-action policy's ability to perform dexterous manipulation tasks on humanoid robots. It specifically probes in-distribution (I.D.) task execution and out-of-distribution (O.O.D.) generalization across varying backgrounds, object placements, and cross-embodiment transfers. Use when the user wants to benchmark on Robot & Human Manipulation Demonstrations, or asks about evaluating this task. Reports Success rate.
Evaluates a language-conditioned transformer model's ability to generate physically plausible and text-aligned 3D humanoid poses from text commands. It probes motion quality, diversity, and multimodal alignment on a retargeted human motion benchmark, as well as real-world deployment success rates. Use when the user wants to benchmark on HumanoidML3D, Humanoid-X, or asks about evaluating this task. Reports FID.
Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
Evaluates the biomechanical plausibility and realism of human motion in AI-generated videos by measuring anatomical, kinematic, and kinetic correctness. It also assesses how well these automated metrics correlate with human preference judgments. Use when the user wants to benchmark on HumanScore Benchmark, or asks about evaluating this task. Reports Kinetic Correctness.