
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the quality of text embedding models in a task-agnostic manner by estimating information sufficiency via normalizing flows. It predicts how well an embedding model will perform on downstream tasks without requiring task-specific labels or fine-tuning. Use when the user has predictions and gold and needs to compute Spearman's ρ.
Evaluates LLMs on complex, multi-step reasoning and agentic search tasks, including single-hop and multi-hop question answering as well as deep research benchmarks requiring web search and synthesis. Use when the user wants to benchmark on NQ, TQA, PopQA, HQA, 2Wiki, MSQ, Bamb, BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
Compute ingyu/klue_mrc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ingyu/klue_mrc.
Evaluates a model's ability to predict traffic accident injury severity (fatal, serious, slight) from tabular accident data. It specifically probes performance on highly imbalanced multi-class classification, focusing on minority-class accuracy and the impact of data imputation and resampling techniques. Use when the user wants to benchmark on UK DfT Traffic Accident Dataset (2005–2019), or asks about evaluating this task. Reports Overall classification accuracy.
Evaluates an AI system's ability to assess scientific research ideas across classification, selection, ranking, and comparison tasks. It probes knowledge-grounded reasoning and multi-perspective evaluation aligned with human expert judgments and conference acceptance standards. Use when the user wants to benchmark on D_point, D_group, D_pair, or asks about evaluating this task. Reports Accuracy.
Evaluates multimodal large language models across general vision, mathematical reasoning, and specialized scientific domains to measure visual perception, instruction following, and domain-specific knowledge retention. Use when the user wants to benchmark on AI2D, OCRBench, ChartQA, MMMU(Val), MMMU-Pro (Standard), MMStar, VStar-Bench, MMBench-EN, MME-RealWorld, DocVQA(Val), InfoVQA(Val), SEED-Bench, SEED-Bench-2-plus, RealWorldQA, MathVision, MathVerse, MathVista, WeMath, ScienceQA, RxnBench,...
Evaluates AI agents' ability to conduct end-to-end LLM research across six domains: data construction, filtering, augmentation, loss/reward design, and scaffold construction. It probes long-horizon decision making, algorithmic robustness, resource management, and iterative code generation in a simulated research environment. Use when the user wants to benchmark on InnovatorBench, or asks about evaluating this task. Reports Best Score.
Evaluates text-to-image retrieval capabilities of vision-language models on expert-level, ecologically grounded queries. It probes fine-grained visual understanding, domain-specific language comprehension, and ranking quality across multiple relevant images per query. Use when the user wants to benchmark on INQUIRE, or asks about evaluating this task. Reports AP@k (mAP@50).
Evaluates multi-class bioacoustic classification of insect audio recordings into one of 459 species. It probes a model's robustness to severe class imbalance, highly variable sampling rates, and ultrasonic frequency ranges. Use when the user wants to benchmark on InsectSet459, or asks about evaluating this task. Reports F1 score.
Evaluates multimodal reasoning and generalized visual search capabilities, measuring how well models can locate and reason about high-information-density images using a multi-agent framework with a dedicated visual search agent. Use when the user wants to benchmark on V*-Bench, Tree-Bench, VisualProbe-Hard, HR-Bench, MME-RealWorld, O3-Bench, or asks about evaluating this task. Reports Accuracy.
Evaluates an LLM's ability to retrieve relevant prior research papers (inspirations) that can inform a given research question from a candidate pool. It measures how well models can surface novel, non-obvious knowledge links through iterative group-based selection. Use when the user wants to benchmark on ResearchBench Inspiration Retrieval, or asks about evaluating this task. Reports Hit Ratio.
Evaluates a model's ability to perform fine-grained, instance-level understanding on images and videos. It probes spatial-temporal grounding, multi-level annotation comprehension (captions, temporal changes), and multiple-choice question answering over explicitly prompted visual regions. Use when the user wants to benchmark on Inst-IT Bench, or asks about evaluating this task. Reports average score.
Evaluates the ability of different instance attribution methods to rank training data instances by their influence on a given test prediction, particularly focusing on identifying problematic training artifacts and comparing gradient-based versus similarity-based approaches. Use when the user wants to benchmark on SST-2, MNLI, HANS, or asks about evaluating this task. Reports Spearman Correlation.
Evaluates a text-to-speech system's ability to follow complex natural-language instructions for acoustic parameter specification, descriptive style direction, and role-play scenarios. It probes fine-grained prosodic control, open-ended style inference, and high-level scenario-based emotional/character expression. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports accuracy.
This evaluation probes a model's ability to generate speech and music conditioned on natural language instructions describing acoustic and musical attributes. It measures text-to-audio fidelity, attribute control accuracy, and perceptual quality across short-form generation tasks. Use when the user wants to benchmark on Seed-TTS benchmark, InstructAudio internal test set, or asks about evaluating this task. Reports WER.
Evaluates instruction-following alignment, truthfulness, toxicity, and bias in large language models. It measures how well model outputs match human preferences and public benchmark standards compared to base models. Use when the user wants to benchmark on API Prompt Distribution, TruthfulQA, RealToxicityPrompts, Winogender, CrowS-Pairs, or asks about evaluating this task. Reports winrate.
This benchmark probes instruction-following capabilities by testing models on 20 carefully designed prompts that enforce format compliance, content constraints, logical sequencing, and multi-step execution. It measures whether models can adhere to verifiable, unambiguous constraints rather than relying on superficial pattern matching or memorized benchmark performance. Use when the user wants to benchmark on Instruction Adherence Diagnostic Prompts, or asks about evaluating this task. Reports...
Evaluates instruction-guided image editing models on their ability to modify an input image according to a text prompt. It measures semantic alignment with the prompt and visual fidelity to a ground-truth edit, plus human preference in pairwise comparisons. Use when the user wants to benchmark on MagicBrush (MagBr), ZONE, or asks about evaluating this task. Reports CLIP-T.
Evaluates whether neural machine translation models can follow diverse natural language instructions (e.g., formality, voice, casing, simplification) without task-specific retraining, while maintaining general translation quality. Use when the user wants to benchmark on WMT'20 News Translation (EN-DE), Multi-30K, Custom Instruction Dataset, or asks about evaluating this task. Reports RR (%).
Evaluates the generalization and domain-adaptive capabilities of language models pre-trained with instruction-augmented corpora. Probes zero/few-shot instruction following, general knowledge, and specialized performance in biomedicine and finance. Use when the user wants to benchmark on MMLU, PubMedQA, ChemProt, RCT, MQP, UMSLE, ConvFinQA, Headline, FiQA SA, FPB, NER, or asks about evaluating this task. Reports average task score.
Evaluates the zero-shot robustness of instruction-tuned language models to variations in instruction phrasing, even when instructions are semantically equivalent. It measures how well models maintain performance on unobserved instruction variants compared to observed ones. Use when the user wants to benchmark on MMLU, BBL, or asks about evaluating this task. Reports accuracy.
Evaluates the instruction-following capability and alignment (helpfulness, honesty, harmlessness) of instruction-tuned LLMs on unseen tasks across English and Chinese. Use when the user wants to benchmark on User-Oriented-Instructions-252, Vicuna-Instructions-80, Unnatural Instructions, or asks about evaluating this task. Reports Relative Score (GPT-4).
Evaluates whether information retrieval models can accurately follow instance-specific, user-aligned instructions rather than generic task descriptions. It probes the robustness of retrievers to instruction variations and their ability to adapt to real-world search scenarios with diverse user contexts. Use when the user wants to benchmark on InstructIR, or asks about evaluating this task. Reports nDCG@10.
Evaluates Vision-Language Models' ability to perform fine-grained visual grounding and instruction reasoning for part segmentation. It probes whether models can infer task-relevant object parts from natural language instructions or oracle prompts, and assesses their capacity for affordance learning in human-robot interaction contexts. Use when the user wants to benchmark on InstructPart, or asks about evaluating this task. Reports gIoU.
Evaluates a text-to-speech model's capability to generate realistic vocal timbres that accurately follow complex natural-language style instructions. It probes fine-grained acoustic control, generalization to unstructured descriptions, and contextual role-play inference. Use when the user wants to benchmark on InstructTTSEval, or asks about evaluating this task. Reports Instruction-following accuracy (%).
Evaluates a text-to-speech model's ability to follow natural-language instructions for voice design, specifically controlling acoustic parameters, descriptive styles, and role-play characteristics. It measures how accurately synthesized speech adheres to explicit semantic and stylistic requirements. Use when the user wants to benchmark on InstructTTSEval-Zh, or asks about evaluating this task. Reports AVG.
Probes phase-aware compliance verification and phase boundary detection in insurance benefit verification calls. It measures a model’s ability to accurately segment conversational phases under workflow-specific rules and apply rule-based compliance reasoning (Information and Procedural Compliance) to fixed spans. Use when the user wants to benchmark on INSURE-Dial, or asks about evaluating this task. Reports exact match (EM).
Evaluates the classification and detection accuracy, as well as inference latency, of neural networks quantized to 8-bit integer arithmetic on mobile ARM CPUs compared to floating-point baselines. Use when the user wants to benchmark on ImageNet, COCO, Face detection dataset, Face attributes dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the calibration and predictive accuracy of ensemble forecasting methods for time-to-event outcomes in meteorology. It compares how well different combination techniques predict the timing of events like the first hard freeze. Use when the user has predictions and gold and needs to compute Mean Integrated Brier Score (IBS).
Evaluates the efficiency of local LLM inference by combining task accuracy with energy consumption to compute Intelligence per Watt (IPW). It probes how model architecture, hardware acceleration, and numerical precision affect the trade-off between performance and power usage on real-world chat and reasoning tasks. Use when the user wants to benchmark on WildChat, NaturalReasoning, SuperGPQA, MMLU Pro, or asks about evaluating this task. Reports accuracy, intelligence per watt (IPW).
This benchmark evaluates the ability of pretrained sentence encoders and classifiers to correctly identify user intents from conversational utterances. It specifically probes few-shot generalization by testing models on severely limited training data (10 or 30 examples per intent) while maintaining a standard full test set. Use when the user wants to benchmark on BANKING77, CLINC150, HWU64, or asks about evaluating this task. Reports accuracy.
Evaluates inter-rater variability among pathologists annotating histopathology images and measures how annotator conformity (agreement with an anchor) impacts downstream deep learning cell detection performance. Use when the user wants to benchmark on Histopathology Cell Annotation Dataset, or asks about evaluating this task. Reports mF1-score.
Evaluates the ability of generative models to synthesize physically plausible, contact-consistent 3D human-object interaction sequences conditioned on text, actions, or object shapes. It probes motion realism, contact accuracy, and alignment between linguistic/action prompts and generated kinematics. Use when the user wants to benchmark on InterAct, or asks about evaluating this task. Reports FID.
Evaluates multimodal large language models' ability to generate functional interactive webpage code from interactive prototype screenshots. It specifically probes the model's capacity to capture dynamic interaction behaviors, element positioning, and visual-textual alignment, rather than just static layout reproduction. Use when the user wants to benchmark on Interaction2Code, or asks about evaluating this task. Reports CLIP.
Evaluates Large Audio Models (LAMs) on real-world, task-oriented voice assistant interactions by capturing user preferences through open-ended pairwise comparisons. It measures how well models align with actual user needs and preferences in an interactive setting, rather than relying on static reference-based benchmarks. Use when the user wants to benchmark on TalkArena Interactive User Preferences, or asks about evaluating this task. Reports Bradley-Terry model score.
Evaluates the effectiveness of interactive document retrieval using user-identified Wikipedia concepts for query expansion and re-ranking. It also tests methods for selecting the most relevant Wikipedia concepts from a large pool based on semantic relevance and document ranking signals. Use when the user wants to benchmark on TREC Filtering-02, HARD-03, HARD-05, or asks about evaluating this task. Reports MAP, P@10.
Evaluates a model's ability to edit two-person 3D motions according to text instructions, balancing semantic modification (instruction adherence) with content preservation (source fidelity) and motion realism. Use when the user wants to benchmark on InterEdit3D, or asks about evaluating this task. Reports Recall@1.
Evaluates the zero-shot sim-to-real transfer and generalization of a Vision-Language-Action (VLA) policy on diverse real-world and simulated manipulation tasks. It probes fundamental pick-and-place, articulated object manipulation, human-robot interaction, and long-horizon task composition capabilities. Use when the user wants to benchmark on InternData-A1 Real-World & Sim-to-Real Benchmarks, or asks about evaluating this task. Reports average success rate.
Evaluates multimodal large language models across general understanding, complex reasoning, mathematics, OCR, document comprehension, and agentic/GUI interaction tasks. Use when the user wants to benchmark on MMMU, MathVista, MMStar, MMVet, or asks about evaluating this task. Reports accuracy.
Evaluates the runtime performance and memory efficiency of a speculatively staged Python interpreter against standard baselines like CPython and PyPy. It probes the interpreter's ability to eliminate dynamic type-checking overhead and optimize instruction dispatch through compile-time specialization. Use when the user wants to benchmark on Computer Language Benchmarks Game, or asks about evaluating this task. Reports speedup.
Evaluates reinforcement learning agents' ability to navigate complex, un-signalized urban intersections under varying traffic conditions. It probes decision-making, collision avoidance, and route completion in dynamic environments with interacting social vehicles. Use when the user wants to benchmark on Intersection Scenarios (RL-CIS), or asks about evaluating this task. Reports Success rate(%).
Evaluates LLM fairness and consistency across intersectional identity attributes (race, gender, socio-economic status) in both ambiguous and disambiguated contexts. It measures accuracy, stereotype alignment, subgroup disparity, and response stability across repeated runs. Use when the user wants to benchmark on Race_SES, Race_Gender, or asks about evaluating this task. Reports Accuracy.
Evaluates the fidelity of natural language explanations for sparse autoencoder (SAE) features by measuring how well an explanation predicts the downstream effects of directly intervening on the feature's activation, rather than just correlating with input contexts. Use when the user has predictions and gold and needs to compute intervention_scoring.
Tests a framework's ability to compress pairwise preference data into interpretable natural language principles (constitutions) and use them to reconstruct original annotations. It probes the model's adaptability to aligned, unaligned, individual, and demographic group preferences, as well as its capacity for bias detection. Use when the user wants to benchmark on Synthetic data, AlpacaEval, Chatbot Arena Conversations, PRISM, or asks about evaluating this task. Reports agreement.
Evaluates a model's capability to perform novel view synthesis, decompose scene properties (albedo, normals, roughness), and relight scenes under new lighting conditions using Gaussian surfels. It specifically probes the model's ability to model indirect illumination and inter-reflections without relying on pre-trained novel view synthesis data. Use when the user wants to benchmark on TensoIR*, Synthetic4Relight*, or asks about evaluating this task. Reports PSNR.
Evaluates the sequential financial decision-making capabilities of LLM-based agents across stock, cryptocurrency, and ETF trading environments. It probes the model's ability to process multi-modal market data, manage portfolio risk, and adapt to volatile market conditions over time. Use when the user wants to benchmark on INVESTORBENCH, or asks about evaluating this task. Reports SR (Sharpe Ratio).
Evaluates LLMs' ability to solve advanced astronomy and astrophysics problems, focusing on geometric/spatial reasoning, physical calculations, and multimodal data analysis. It benchmarks performance against human Olympiad participants using official scoring rubrics. Use when the user wants to benchmark on IOAA (International Olympiad on Astronomy and Astrophysics), or asks about evaluating this task. Reports score.
This evaluation probes an intrusion detection model's ability to classify network traffic flows as benign or malicious across highly imbalanced IoT datasets. It specifically tests the model's robustness to extreme class imbalance and its capacity to leverage graph-structured representations of network flows for anomaly detection. Use when the user wants to benchmark on BoT-IoT, ToN-IoT, or asks about evaluating this task. Reports macro-F1.
Evaluates the ability of a lightweight CNN to classify IoT binary files as benign or belonging to specific DDoS malware families (Mirai, Linux.Gafgyt) by converting raw binaries into 64x64 grayscale images. Use when the user wants to benchmark on IoTPOT IoT DDoS Malware Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of a graph neural network to detect network intrusions in IoT environments by classifying traffic flows as benign or malicious, and identifying specific attack types. It probes the model's capacity to leverage topological graph structures and edge features for robust intrusion detection across imbalanced, real-world network traffic datasets. Use when the user wants to benchmark on BoT-IoT, NF-BoT-IoT, ToN-IoT, NF-ToN-IoT, or asks about evaluating this task. Reports F1-Sc...