
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark evaluates the human-centric video understanding capabilities of multimodal large language models (MLLMs). It specifically probes inner emotion perception, outer behavioral manifestations, and cross-modal speech-visual alignment through 16 fine-grained multiple-choice tasks. Use when the user wants to benchmark on HumanVBench, or asks about evaluating this task. Reports accuracy.
Probes the generalizability, diversity, and reconstruction accuracy of 3D human body expression models (gaze, face, hand, body) across multiple datasets and viewpoints. It evaluates how well models trained on HUMBI generalize to unseen datasets and how accurately they reconstruct 3D geometry from monocular images. Use when the user wants to benchmark on HUMBI, or asks about evaluating this task. Reports IoU.
Evaluates the emotional intelligence of audio language models across multi-turn dialogues. It probes four core capabilities: tracking emotional trajectories over time, reasoning about implicit emotional causes, generating empathetic responses, and resolving conflicts between acoustic and textual emotional signals. Use when the user wants to benchmark on HumDial-EIBench, or asks about evaluating this task. Reports Accuracy (%).
This benchmark evaluates human-like spoken dialogue systems across two core capabilities: emotional intelligence (multi-turn emotion tracking, causal reasoning, and empathetic response generation) and full-duplex interaction (natural turn-taking, interruption handling, and noise rejection during concurrent listening and speaking). It uses authentic real-world conversations to measure long-term emotional consistency and cognitive synchronization. Use when the user wants to benchmark on HumDial...
Evaluates text embedding models against human baselines across 16 MTEB datasets, probing semantic similarity, classification, clustering, and reranking capabilities. It specifically measures cross-lingual performance and identifies task ambiguities where model scores may reflect label pattern reproduction rather than genuine understanding. Use when the user wants to benchmark on MTEB (16 datasets, 26 task-language pairs), or asks about evaluating this task. Reports accuracy.
Evaluates a humanoid robot's capability to learn and execute whole-body manipulation skills from robot-free demonstrations, assessing manipulation precision, dynamic coordination, generalization to unseen environments/objects, and data-collection efficiency. Use when the user wants to benchmark on Squatting, Tossing, Bimanual Actions, Walking, Dynamic Motions, or asks about evaluating this task. Reports success rate.
Evaluates large audio-language models on music understanding by testing their ability to answer multiple-choice questions about audio excerpts. It probes perceptual grounding, structural/harmonic/cultural reasoning, and robustness against text-only shortcuts or answer-position bias. Use when the user wants to benchmark on HumMusQA, or asks about evaluating this task. Reports accuracy.
Evaluates LLM humor generation by conducting pairwise preference judgments on joke outputs across 300 diverse prompts. It measures how well models master comedic mechanisms rather than relying on model scale, using an LLM-as-judge framework to produce global rankings. Use when the user wants to benchmark on SemEval-2026 MWAHAHA, or asks about evaluating this task. Reports Bradley-Terry Maximum Likelihood Estimation.
Evaluates NLP models on patent-related tasks including binary classification of patent acceptance, multi-class subject area classification using IPC codes, and abstractive summarization of patent claims or descriptions into abstracts. Use when the user wants to benchmark on Harvard USPTO Patent Dataset (HUPD), or asks about evaluating this task. Reports accuracy.
Evaluates human-centric video anomaly detection models on continuously recorded real-world video streams. It measures how well models distinguish between normal and anomalous human activities across multiple camera views, particularly under continual learning conditions. Use when the user wants to benchmark on HuVAD, or asks about evaluating this task. Reports AUC-ROC.
This benchmark evaluates a model's ability to perform high-order, multistep visual question answering by integrating visual scene graphs with external commonsense knowledge. It explicitly probes the model's reasoning process by requiring it to predict intermediate knowledge triplets alongside the final answer, enforcing explainability and self-diagnosis capabilities. Use when the user wants to benchmark on HVQR, or asks about evaluating this task. Reports triplet recall.
Evaluates hardware-aware neural architecture search (HW-NAS) algorithms by measuring how effectively they discover network topologies that optimize the trade-off between classification accuracy and on-device inference latency for specific target hardware. Use when the user wants to benchmark on HW-NAS-Bench, or asks about evaluating this task. Reports top-1 accuracy.
Evaluates the capability of DNA foundation models to perform short-range and long-range genomic understanding tasks, as well as their ability to generate biologically plausible cis-regulatory elements. It probes sequence classification, variant effect prediction, and generative design across multiple species and cell types. Use when the user wants to benchmark on GUE, BEND, LRB, CRE (regLM), or asks about evaluating this task. Reports MCC.
Multi-hop question answering that requires integrating information from both tabular and textual sources. It probes a model's ability to perform cross-modal reasoning and extract precise answers from heterogeneous data. Use when the user wants to benchmark on HybridQA, or asks about evaluating this task. Reports exact match (EM).
Evaluates retrieval-augmented models' ability to perform multi-hop reasoning over hybrid knowledge (unstructured text and knowledge graphs) using time-framed, external scientific literature to prevent parametric memorization. Use when the user wants to benchmark on Arxiv-AI, Arxiv-CY, Arxiv-BIO, or asks about evaluating this task. Reports accuracy.
Evaluates semi-supervised and traditional machine learning models for detecting hydraulic system anomalies using only normal data for training. It probes the ability of models to generalize from normal-condition features and identify leakage faults under class-imbalanced testing conditions. Use when the user wants to benchmark on Unspecified hydraulic condition monitoring dataset, or asks about evaluating this task. Reports F1_Score.
This protocol evaluates face-based voice conversion models by measuring how well synthesized audio matches the target speaker's identity and pitch characteristics using only facial images as input. It probes cross-modal alignment, speaker homogeneity, diversity, and explicit fundamental frequency (F0) estimation accuracy. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports Pitch deviation.
Compute hynky/sklearn_proxy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hynky/sklearn_proxy.
Evaluates hyperspectral foundation models and task-specific deep learning architectures on pixel-level regression tasks for retrieving cloud optical and microphysical properties (COT, CER, CWP, CTH) from NASA PACE-OCI imagery. Use when the user wants to benchmark on HyperFM250K, or asks about evaluating this task. Reports MSE.
Evaluates a multimodal model's ability to generate and anticipate scene graphs from video frames, capturing spatial object relationships and causal temporal transitions. It also tests video question answering, captioning, and relation reasoning capabilities by leveraging hypergraph structures to model multi-way interactions. Use when the user wants to benchmark on VSGR, PVSG, Action Genome, or asks about evaluating this task. Reports Recall (R) / mean Recall (mR).
Evaluates the ability of mRNA language models to predict diverse biological properties (e.g., protein expression, degradation, thermostability) and annotate antibody sequence regions. It also probes model robustness to out-of-distribution sequence lengths and extreme GC content, testing generalization in hierarchical biological representation learning. Use when the user wants to benchmark on Ab1, Ab2, mRFP, COVID-19 Vaccine, Drosophila melanogaster, Saccharomyces cerevisiae, Pichia pastoris, ...
Evaluates the optimization quality and time efficiency of hyperparameter search algorithms by comparing the test error rate of recommended configurations against wall-clock time across neural architecture and traditional ML benchmarks. It measures how quickly each optimizer converges to near-optimal configurations under sequential and parallel deployment settings. Use when the user wants to benchmark on NATS-Bench, LIBSVM Covertype, or asks about evaluating this task. Reports test_error_rate.
Compute hyperml/balanced_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hyperml/balanced_accuracy.
Evaluates the accuracy and uncertainty quantification of a physics-informed neural network for locating earthquake hypocenters using synthetic seismic arrival times. Use when the user wants to benchmark on Synthetic Seismic Array, or asks about evaluating this task. Reports location uncertainty.
Assesses an LLM's capability to synthesize novel research hypotheses by combining a given research background with retrieved inspiration papers. It probes the model's ability to mutate and recombine scientific concepts into coherent, groundtruth-aligned proposals. Use when the user wants to benchmark on ResearchBench Hypothesis Composition, or asks about evaluating this task. Reports Normalized Composition Score.
Measures an LLM's ability to correctly rank a groundtruth hypothesis against a set of negative hypotheses using pairwise comparisons. It evaluates discriminative judgment in scientific reasoning. Use when the user wants to benchmark on ResearchBench Hypothesis Ranking, or asks about evaluating this task. Reports Accuracy.
Evaluates the retrieval capability of hybrid (dense + sparse/lexicon) models on Chinese text. It probes how well the model ranks relevant passages for a given query across diverse Chinese domains like medical, e-commerce, and general web. Use when the user wants to benchmark on C-MTEB, or asks about evaluating this task. Reports nDCG@10.
This evaluation probes how anisotropic regularization (I-STAR) affects the downstream performance of fine-tuned language models across standard NLP benchmarks. It also measures the geometric properties of the resulting embedding spaces, specifically isotropy and intrinsic dimensionality, to correlate representation structure with task accuracy. Use when the user wants to benchmark on SST-2, QNLI, RTE, MRPC, QQP, COLA, STS-B, SST-5, SQUAD, or asks about evaluating this task. Reports accuracy.
Evaluates the classification performance of Spiking Neural Networks trained on synthetic event streams generated from static images, and tests the transferability of these models to real-world neuromorphic sensor data. Use when the user wants to benchmark on I2E-CIFAR10, I2E-CIFAR100, I2E-ImageNet, CIFAR10-DVS, or asks about evaluating this task. Reports Accuracy.
Evaluates LLMs' ability to perform structured reasoning and knowledge application across arithmetic, logical, commonsense, and symbolic tasks using a template-based prompting framework. Use when the user wants to benchmark on GSM8K, AQuA, Date Understanding, Object Tracking, StrategyQA, CommonsenseQA, Last Letter, or asks about evaluating this task. Reports accuracy.
Evaluates machine learning models' ability to classify electron antineutrino (IBD) events from background accidents in a liquid scintillator detector. It measures how well the models preserve signal efficiency while controlling background contamination compared to traditional cut-based selection. Use when the user wants to benchmark on JUNO IBD/Accident Dataset, or asks about evaluating this task. Reports efficiency.
Evaluates whether machine learning models can correctly capture non-additive interaction effects in drug-target affinity prediction, rather than merely learning global means or individual drug/target main effects. It measures the proportion of correctly predicted interaction directions across test pairs. Use when the user has predictions and gold and needs to compute IC-index.
Evaluates computational argumentation solvers on their ability to correctly compute extensions (e.g., semi-stable, stage, ideal) across diverse argumentation frameworks ranging from random graphs to application-derived structures. Use when the user wants to benchmark on ICCMA'17 Benchmark Suite, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates end-to-end document understanding on low-quality scanned receipts, specifically testing text localization, character-level OCR, and structured key information extraction (e.g., company, cash, date, address). Use when the user wants to benchmark on ICDAR2019 SROIE, or asks about evaluating this task. Reports primary metric.
Evaluates image generation and editing models across 31 fine-grained tasks spanning text-to-image creation, reference-guided creation, and various editing scenarios. It probes capabilities in aesthetic quality, imaging quality, prompt adherence, source/reference consistency, and controllability. Use when the user wants to benchmark on ICE-Bench, or asks about evaluating this task. Reports prompt following (PF).
Evaluates bilingual (Chinese and English) financial large language models across 14 NLP tasks, including sentiment analysis, classification, question answering, and information extraction. It probes cross-lingual adaptability, domain-specific reasoning, and instruction-following capabilities on financial text. Use when the user wants to benchmark on FE, StockB, CFPB, CFiQA-SA, FPB, FiQA-SA, Corpus, AFQMC, NL, NL2, NSP, FinevalF, StcokA, CACL18, CBigData18, CIKM18, ACL18, BigData18, RE, CHeadl...
Measures intervention consistency in LLM decision-making by checking whether swapping irrelevant features (demographic names, authority credentials, or framing phrasing) causes the model to change its verdict. Probes susceptibility to spurious feature reliance and systematic bias across high-stakes domains. Use when the user wants to benchmark on ICE-Guard Benchmark, or asks about evaluating this task. Reports flip_rate.
Evaluates the cross-cohort generalisability of transcriptomic models (bulk and single-cell RNA-seq) for predicting immune checkpoint inhibitor (ICI) response in cancer patients. Probes robustness to cohort-specific transcriptomic context, tumour type, immune composition, and class imbalance. Use when the user wants to benchmark on Cho et al., Ribas et al., Poddubskaya et al., Gondal et al., Franken et al., Luoma et al., Reinstein et al., or asks about evaluating this task. Reports macro F1 sc...
Evaluates embedding models and rerankers on their ability to retrieve contextually useful documents for In-Context Learning (ICL) tasks. It measures retrieval effectiveness by ranking candidate documents based on their utility in improving downstream LLM accuracy, rather than relying solely on semantic similarity. Use when the user wants to benchmark on TruthfulQA, Emotion, ProductER, or asks about evaluating this task. Reports nDCG@10.
Probes the causal impact of LLM-assisted peer reviews on paper scoring and acceptance outcomes at a major machine learning conference. It measures whether AI-assisted reviews systematically inflate scores and increase acceptance probabilities, particularly for borderline submissions. Use when the user wants to benchmark on ICLR Conference Reviews (2018-2024), or asks about evaluating this task. Reports acceptance_rate_difference.
Quantifies the average research effort required to produce one publication at top-tier conferences across 27 computer science subfields. It enables cross-area comparisons of faculty productivity and publication effort by normalizing faculty headcounts against publication counts. Use when the user has predictions and gold and needs to compute ICLR points.
Evaluates the causal impact of LLM-generated feedback on peer review quality by measuring reviewer engagement, revision rates, feedback incorporation, and downstream rebuttal dynamics in a large-scale randomized controlled trial. Use when the user wants to benchmark on ICLR 2025 Reviews, or asks about evaluating this task. Reports update_rate.
Evaluates whether text-based LLMs can effectively leverage high-dimensional non-text modality representations (e.g., molecular embeddings from foundation models) via training-free in-context learning, comparing various representation injection and projection strategies. Use when the user wants to benchmark on ESOL, Caco_wang, AqSolDB, LD50_Zhu, AstraZeneca, or asks about evaluating this task. Reports RMSE.
Evaluates machine learning models' ability to detect and classify cyberattacks in Industrial Control Systems (ICS) network traffic. It probes the capability to distinguish between normal operations and specific attack types like DDoS, IP-Scan, MitM, Port-Scan, and Replay using flow-level features. Use when the user wants to benchmark on ICS-Flow, or asks about evaluating this task. Reports F1-score.
Predicts the risk of a patient being readmitted to the ICU within 30 days of discharge using longitudinal electronic medical record (EMR) data. The task evaluates how well different deep learning architectures can model time-varying clinical events (diagnoses, procedures, medications, vital signs) alongside static demographic covariates to capture complex patient trajectories and risk factors. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.
Evaluates the robustness and safety of semantic segmentation models for autonomous driving in unstructured traffic and adverse weather. It specifically probes whether models can correctly identify critical road elements and traffic participants when visual quality degrades due to rain, fog, snow, or low light. Use when the user wants to benchmark on IDD-AW, or asks about evaluating this task. Reports Safe mIoU (SmIoU).
Evaluates a pipeline's ability to transform vague research intents into structured, methodologically grounded research patterns. It probes the system's capacity for methodological abstraction, knowledge graph retrieval, and coherent scientific narrative generation compared to direct LLM prompting. Use when the user wants to benchmark on ICLR & NeurIPS Papers (3-year corpus), or asks about evaluating this task. Reports LLM-judged preference (novelty, methodological substance, overall research ...
Evaluates the professional design capabilities of generative models across text-to-image, image-to-image, and multi-image generation tasks. It probes aesthetic quality, contextual relevance, multimodal alignment, and adherence to complex, real-world design requirements that go beyond basic generation. Use when the user wants to benchmark on IDEA-Bench, or asks about evaluating this task. Reports Avg. Score.
Evaluates a framework's ability to decompose scientific papers into orthogonal conceptual dimensions (problem, method, findings) and model transitions between them. It probes fine-grained conceptual similarity retrieval and assesses whether the model's novelty predictions align with expert human judgments. Use when the user wants to benchmark on ICLR 2025 Submissions, AI-Researcher, or asks about evaluating this task. Reports Recall@K.
This evaluation probes a dialogue system's ability to dynamically generate derived questions and manage multi-turn interactions to accurately classify loan applicants as fraudulent or legitimate based on their knowledge of personal information triplets. Use when the user wants to benchmark on Applicant Personal Information Dataset, or asks about evaluating this task. Reports recognition accuracy.