
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multimodal LLMs' ability to perform spatiotemporal reasoning and probabilistic risk forecasting for tornadoes by interactively querying weather data and generating geographic risk polygons. It measures forecasting accuracy, hallucination severity, and geometric precision against official meteorological baselines. Use when the user wants to benchmark on TornadoBench, or asks about evaluating this task. Reports TornadoBench.
Evaluates large language models' context-sensitive reasoning and decision-making capabilities in autonomous driving scenarios. It probes physics-based calculations, policy compliance, risk interpretation, and maneuver optimization through multiple-choice questions derived from structured driving simulations. Use when the user wants to benchmark on AgentDrive-MCQ, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates AI agents and human-AI collaboration on domain-specific data science tasks across six industries. It probes the ability to perform feature engineering, integrate multimodal data (images, text, PDFs, JSON), and build predictive models that require genuine domain reasoning rather than generic pipelines. Use when the user wants to benchmark on AgentDS, or asks about evaluating this task. Reports quantile_score.
Evaluates autonomous clinical decision-making agents on Electronic Health Record (EHR) data. It probes multi-step reasoning, long-context dependency preservation, and robustness to distribution shifts across different hospital databases and clinical event types. Use when the user wants to benchmark on MIMIC-IV / MIMIC-III, or asks about evaluating this task. Reports average score.
Evaluates LLM-based data analysis agents on their ability to execute domain-specific time-series queries, particularly focusing on stateful logic, temporal dependencies, and incident pattern detection. It probes whether agents can correctly interpret schemas, track state across sequential events, and identify anomalous behavior without relying on predefined time windows. Use when the user wants to benchmark on AgentFuel Benchmark, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the harmfulness and safety alignment of LLM-based agents by measuring their compliance with malicious, multi-step tasks that require coherent tool chaining. It probes whether models can be coerced into executing harmful behaviors through direct prompting or simple jailbreak templates, while tracking refusal rates and capability preservation. Use when the user wants to benchmark on AgentHarm, or asks about evaluating this task. Reports harm score.
Evaluates the ability of embodied multi-agent systems to execute long-horizon, coordinated tasks efficiently using cache-driven asynchronous planning. It probes how well agents can reuse cached plan transitions to reduce LLM inference latency and token costs while maintaining high task success rates across diverse 3D simulation environments. Use when the user wants to benchmark on TDW-MAT, TDW-COOK, TDW-GAME, BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate.
Evaluates LLM agents' ability to navigate simulated environments and execute multi-step plans to complete natural language instructions. It probes step-wise decision-making, goal proximity tracking, and sequential task execution across web shopping, grid-world navigation, and text-based crafting scenarios. Use when the user wants to benchmark on WebShop, BabyAI, TextCraft, or asks about evaluating this task. Reports success rate.
This evaluation protocol measures LLM agent performance on multi-step reasoning tasks by tracking step-wise progress toward goal completion and the frequency of repetitive actions or states. It enables fine-grained debugging and architectural refinement beyond simple pass/fail success rates. Use when the user wants to benchmark on ALFWorld, Sudoku, or asks about evaluating this task. Reports progress rate.
This benchmark evaluates LLM-based agentic recommender systems across three scenarios: classic, evolving-interest, and cold-start recommendation. It probes the agents' ability to dynamically plan, utilize textual interaction environments, and adapt to user preference shifts or data sparsity using structured user/item profiles and reviews. Use when the user wants to benchmark on Amazon, GoodReads, Yelp, or asks about evaluating this task. Reports Hit Rate@$N.
This benchmark evaluates the effectiveness of LLM-based judges in automatically assessing web agent trajectories. It probes the judges' ability to correctly predict task success, detect side effects, and identify repetitive actions by comparing their outputs against expert human annotations. Use when the user wants to benchmark on AgentRewardBench, or asks about evaluating this task. Reports precision.
Evaluates the safety of embodied vision-language model agents when executing hazardous instructions in simulated indoor environments. It probes four key capabilities across perception, planning, and execution stages: object recognition accuracy, refusal to plan harmful actions, success in generating harmful plans, and success in physically executing them under both direct and jailbroken conditions. Use when the user wants to benchmark on AGENTSAFE, or asks about evaluating this task. Reports ...
Probes the clinical faithfulness, factual accuracy, and diagnostic logic of medical imaging report generation systems. It evaluates robustness to paraphrasing and semantic perturbations by decomposing assessment into interpretable reasoning stages that mimic radiologist workflows. Use when the user wants to benchmark on Five medical imaging datasets (names not provided in excerpt), or asks about evaluating this task. Reports AgentsEval Score.
Evaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.
Evaluates the ability of multimodal agents to perform long-horizon, multi-step tool use in complex, realistic visual environments. It probes cross-image reasoning, constraint tracking, and robust grounding when interacting with dynamic tools like web search and code execution. Use when the user wants to benchmark on AgentVista, or asks about evaluating this task. Reports accuracy.
Evaluates the perceptual quality and text-image correspondence of AI-generated human images, while also benchmarking the ability of models to identify visible and semantically distorted human body parts. Use when the user wants to benchmark on AGHI-QA, or asks about evaluating this task. Reports SRCC.
This benchmark evaluates foundation models on human-level cognitive abilities and general reasoning by testing them on a diverse collection of standardized admission and qualification exams. It probes domain-specific knowledge, analytical reasoning, and problem-solving across subjects like mathematics, law, logic, and languages. Use when the user wants to benchmark on AGIEval, or asks about evaluating this task. Reports accuracy.
Compute agkphysics/ccc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of agkphysics/ccc.
Evaluates semi-supervised node classification on citation networks using an attention-based graph neural network. Tests performance under fixed benchmark splits, random node sampling, and larger training sets to measure classification accuracy and attention interpretability. Use when the user wants to benchmark on CiteSeer, Cora, PubMed, or asks about evaluating this task. Reports classification accuracy.
Evaluates the ability of LLMs to generate accurate and context-aware agricultural recommendations (sowing schedules, irrigation plans, risk mitigation) based on integrated weather, soil, and crop data. It specifically probes how multi-round prompt engineering improves recommendation quality compared to single-round and Chain-of-Thought baselines. Use when the user wants to benchmark on Agricultural Meteorological Dataset, or asks about evaluating this task. Reports Accuracy (Acc).
Evaluates Automatic Speech Recognition (ASR) models on real-world agricultural field recordings across three Indian languages (Hindi, Telugu, Odia). It probes the models' ability to transcribe domain-specific terminology under challenging acoustic conditions like wind noise and multi-speaker overlap. Use when the user wants to benchmark on Agricultural Field Recordings, or asks about evaluating this task. Reports AWWER.
Evaluates a unified speech-vision-text model's capability in multilingual agricultural reasoning, covering text generation, vision-language QA, and multimodal speech understanding across open-ended and multiple-choice formats. Use when the user wants to benchmark on AgriBench-13K, AgriBench-VL-4K, AgriBench-Omni-2K, or asks about evaluating this task. Reports Accuracy.
Evaluates an agent's ability to decompose abstract, long-horizon household instructions into feasible, constraint-satisfying action plans. It probes intent inference, subgoal grounding, and robustness to environmental clutter and instruction ambiguity. Use when the user wants to benchmark on AHAT, Human Tasks, PARTNR, Behavior-1K, or asks about evaluating this task. Reports Success Rate (SR).
Compute ahnyeonchan/Alignment-and-Uniformity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ahnyeonchan/Alignment-and-Uniformity.
Evaluates monocular 3D human pose estimation models trained exclusively on synthetic 3D data and real 2D images, testing their ability to generalize to real-world 3D pose benchmarks without using any real 3D pose annotations during training. It probes domain adaptation capabilities, cross-dataset generalization, and the effectiveness of skeletal pose alignment strategies. Use when the user wants to benchmark on Human3.6M, MuPoTS, SURREAL, ScanAva+, MSCOCO, MPII Human Pose, or asks about evalu...
Evaluates the computational performance and energy efficiency of various AI accelerators (CPUs, GPUs, TPUs) across standard deep learning workloads, including CNNs and NLP models. It measures how hardware architecture, numerical precision, and batch size impact training throughput and power consumption. Use when the user wants to benchmark on Standard DNN Workloads (ResNet50, Inception v3, Vgg16, LSTM, Deep Speech 2, Transformer), or asks about evaluating this task. Reports throughput.
Evaluates the inference performance of mobile AI accelerators across major SoC vendors by running a standardized suite of deep learning models via TensorFlow Lite and NNAPI. It measures latency and accuracy to compare on-device AI capabilities against desktop hardware and track hardware evolution. Use when the user wants to benchmark on AI Benchmark 3.0, or asks about evaluating this task. Reports AI-Score.
Evaluates the fairness and utility of AI-generated face detectors across demographic attributes (skin tone, gender, age) and intersectional groups. It measures how well detectors distinguish real vs. AI-generated faces while ensuring equitable performance across demographic subgroups. Use when the user wants to benchmark on AI-Face, or asks about evaluating this task. Reports $F_{MEO}$.
Evaluates the ability of AI-generated image detectors to generalize to novel, temporally subsequent generative models under realistic post-processing conditions. It measures how well detectors maintain performance when incrementally trained on historically ordered synthetic data and tested on unseen future generators. Use when the user wants to benchmark on AI-GenBench, or asks about evaluating this task. Reports AUROC, Accuracy.
Evaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.
Evaluates a feature-hierarchical edge inference framework's ability to dynamically allocate communication and computation resources to maximize AI quality under strict latency and energy constraints. Use when the user has predictions and gold and needs to compute AI quality (mAP).
Evaluates adversarial AI red-teaming capabilities by measuring participant success rates in bypassing LLM guardrails, manipulating model outputs, and extracting sensitive data through prompt injection and jailbreaking techniques. Use when the user wants to benchmark on AI Red Teaming CTF (CTF ID: 2604), or asks about evaluating this task. Reports solve_rate.
Evaluates a style-based classifier's ability to detect AI-generated text in academic peer reviews and measures temporal generalization by tracking detection rates across consecutive years. Use when the user wants to benchmark on ICLR Peer Reviews, Nature Communications Peer Reviews, or asks about evaluating this task. Reports percentage_ai_detected.
Evaluates how source disclosure and perceived AI authorship influence human editing behavior and subsequent peer-review acceptance decisions for scientific abstracts. Use when the user wants to benchmark on CS-Conference-Abstracts, or asks about evaluating this task. Reports accept/reject decision.
This benchmark evaluates pixel-wise sea ice stage of development (SOD) segmentation using dual-polarized SAR imagery. It probes a model's ability to accurately classify ice types under varying quantization levels and measures hardware efficiency across different computing platforms. Use when the user wants to benchmark on AI4Arctic Sea Ice Dataset, or asks about evaluating this task. Reports F1 score.
Evaluates histopathology foundation models' ability to extract center-invariant, biologically relevant features for skin cancer subtyping. It measures representation bias toward scanning centers and downstream classification performance under multiple instance learning frameworks. Use when the user wants to benchmark on AI4SkIN, or asks about evaluating this task. Reports Balanced Accuracy (BACC).
Probes the end-to-end latency and micro-architectural efficiency of AI-accelerated internet service workloads. It measures how AI components impact service latency and GPU execution stalls during both online inference and offline training. Use when the user wants to benchmark on AIBench E-commerce Search Workload, or asks about evaluating this task. Reports Latency (avg, p90, p99).
Evaluates the end-to-end system-level performance and tail latency of AI-driven online services by simulating real-world user workloads. It probes how cascading interactions between AI and non-AI components affect overall service quality, and tests the validity of statistical queueing models for predicting system latency. Use when the user wants to benchmark on AIBench Scenario (E-commerce & Translation Intelligence), or asks about evaluating this task. Reports latency (avg, p90, p99).
Evaluates AI training workloads by measuring model complexity, computational cost, convergence rate, and micro-architectural behavior to assess benchmark diversity, repeatability, and cost against industry standards like MLPerf. Use when the user wants to benchmark on AIBench Training, or asks about evaluating this task. Reports convergent_rate.
Evaluates Vision-Language Models on affective image content analysis across three dimensions: Emotion Understanding (identifying emotions in images), Emotion Reasoning (inferring emotional causes/context), and Emotion-Guided Content Generation (producing text guided by emotional intent). It probes models' ability to perceive, reason about, and generate content based on visual emotional cues, including sensitivity to abstract art and reliance on facial shortcuts. Use when the user wants to ben...
Evaluates the ability of computer vision models to classify aerial imagery into distinct scene categories. It probes robustness to high intra-class diversity and low inter-class similarity in remote sensing data. Use when the user wants to benchmark on AID, or asks about evaluating this task. Reports accuracy.
Evaluates LLM-driven document analysis agents on end-to-end data analytics workflows, including question answering, data visualization, and file generation. It probes multi-step numerical reasoning, cross-data consistency, and long-horizon planning on heterogeneous real-world documents. Use when the user wants to benchmark on AIDABench, or asks about evaluating this task. Reports Pass@3.
Evaluates the performance of LLMs fine-tuned on synthetic data generated by AIDE across a suite of standard knowledge and reasoning benchmarks. It probes zero-shot and few-shot generalization capabilities compared to models fine-tuned on human-curated gold data. Use when the user wants to benchmark on MMLU, FinBen, ARC-Challenge, GSM8K, TruthfulQA, MedQA, BIG-Bench, or asks about evaluating this task. Reports zero-shot accuracy.
Evaluates the effectiveness of AI-generated outpainted vehicle images as data augmentation for training object detection models. It probes the model's ability to generalize to real-world vehicle classification and bounding box localization when trained on synthetically augmented data. Use when the user wants to benchmark on AIDOVECL augmented dataset, or asks about evaluating this task. Reports F1 Score.
Evaluates a model's ability to forecast daily specific streamflow at global gauging stations under temporal generalization. It specifically probes robustness to domain shifts between reanalysis pre-training data and operational forecast fine-tuning data, testing whether the model maintains performance when transitioning from historical reanalysis to real-time operational forcing. Use when the user wants to benchmark on CARAVAN v1.5, or asks about evaluating this task. Reports KGE.
Evaluates the cross-generator generalization capability of AI-generated image (AIGC) detectors. It probes whether models trained on a specific generator (SDv1.4) can accurately distinguish real from fake images produced by diverse, unseen generative models and in-the-wild sources. Use when the user wants to benchmark on GenImage, GenImage++, Chameleon, or asks about evaluating this task. Reports Accuracy (ACC).
Evaluates the performance of image-to-video (I2V) generation models across multiple quality and alignment dimensions. It probes how well models preserve input image fidelity, generate coherent motion, align with text prompts, maintain temporal consistency, and produce high-quality video output. Use when the user wants to benchmark on AIGCBench Dataset, or asks about evaluating this task. Reports video quality.
Evaluates the ability of models to distinguish real photographs from AI-generated images across diverse, out-of-distribution, and post-processed scenarios. It probes both low-level pixel artifact detection and high-level semantic consistency checking to measure real-world generalization. Use when the user wants to benchmark on Chameleon, WildRF, AIGI-Bench, Co-SPY-Bench (in-the-wild), BFree-Online, AIGI-Now, GenImage, DRCT-2M, AIGCDetectBenchmark, or asks about evaluating this task. Reports B...
Evaluates the generalization, robustness to image degradation, and sensitivity to data augmentation and pre-processing of AI-generated image (AIGI) detectors across 25 diverse test datasets spanning GANs, diffusion models, and face-swap/manipulation methods. Use when the user wants to benchmark on AIGIBench, or asks about evaluating this task. Reports F.Acc..
This benchmark evaluates the perceptual quality and text-to-image alignment of AI-generated images. It benchmarks objective quality assessment models against large-scale human subjective ratings to measure how well automated metrics correlate with human perception. Use when the user wants to benchmark on AIGIQA-20K, or asks about evaluating this task. Reports SRoCC.