
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates multimodal large language models' ability to perform online spatio-temporal scene understanding and dynamic, agent-centric reasoning. It tests how well models update spatial and temporal knowledge as they incrementally explore environments, retrieve long-term memory, and infer object relationships across sequential observations. Use when the user wants to benchmark on OST-Bench, or asks about evaluating this task. Reports Overall average score.
Evaluates deep learning models on classifying osteosarcoma histopathology images into non-tumor, non-viable tumor, viable tumor, and non-viable ratio categories without prior segmentation. Probes the model's ability to capture local texture and global spatial patterns for medical image classification. Use when the user wants to benchmark on TCIA Osteosarcoma, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal agents' ability to perform open-ended, real-world desktop computer tasks across multiple operating systems. It probes GUI grounding, multi-app workflow navigation, and executable action prediction in dynamic, interactive environments. Use when the user wants to benchmark on OSWorld, or asks about evaluating this task. Reports success.
Evaluates an agent's ability to autonomously plan and execute multi-step GUI automation tasks across various desktop applications. It measures the success rate on both in-distribution tasks from the OSWorld-Verified benchmark and out-of-distribution tasks across six distinct Linux applications. Use when the user wants to benchmark on OSWorld-Verified, OOD GUI Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates a machine learning model's ability to detect overshooting tops (OTs) in satellite imagery at a 2 km pixel resolution. It measures how well the model predicts convection/OT presence using physics-informed features derived from visible and infrared channels. Use when the user wants to benchmark on GOES-16 ABI + MRMS Convection Labels, or asks about evaluating this task. Reports hit, correct rejection, false alarm, miss counts.
Evaluates the energy efficiency and update latency of Over-The-Air (OTA) firmware update strategies on flash-based, batteryless IoT devices under simulated energy-harvesting conditions. Use when the user wants to benchmark on OTA Firmware Update Benchmarks, or asks about evaluating this task. Reports Total Update Energy Consumption.
Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
This benchmark evaluates a model's ability to perform open-domain question answering by retrieving and fusing evidence from both tabular and textual sources. It specifically probes multi-hop reasoning capabilities where answers require bridging information across separate table segments and text passages. Use when the user wants to benchmark on OTT-QA, or asks about evaluating this task. Reports EM.
Evaluates the ability of an outlier detection algorithm to identify anomalous data points in highly imbalanced datasets without prior knowledge of fraud patterns. It probes consistency estimation and ensemble clustering robustness across varying feature spaces and class distributions. Use when the user wants to benchmark on Satimage-2, Thyroid, Credit Card Fraud Detection, or asks about evaluating this task. Reports AUPRC.
Evaluates a model's ability to localize objects in images based on natural language descriptions without prior exposure to those specific categories. It probes visual-linguistic alignment, handling of novel vocabulary, and robustness to varying object scales and complex scenes. Use when the user wants to benchmark on OV-VG, or asks about evaluating this task. Reports Acc50.
Probes a model's ability to generate open-vocabulary video scene graphs by predicting objects, attributes, relations, and triplets from ground-truth trajectories. It evaluates semantic matching accuracy beyond exact label overlap and tests temporal grounding for dynamic relations. Use when the user wants to benchmark on PVSG, VidOR, VIPSeg, SVG2 test set, or asks about evaluating this task. Reports object/attribute/relation/triplet prediction accuracy (LLM-judged).
Evaluates histopathology foundation models and ImageNet-pretrained encoders on classifying ovarian cancer subtypes from whole slide images. It probes the ability of vision models to extract diagnostically relevant features from medical histology slides for multi-class classification. Use when the user wants to benchmark on Ovarian Cancer WSI Dataset, or asks about evaluating this task. Reports balanced accuracy.
This benchmark probes a model's ability to recognize visual entities in an open-domain setting, specifically testing zero-shot generalization to entities not seen during training. It evaluates how well a model can align visual inputs with structured knowledge graph descriptions to perform entity retrieval. Use when the user wants to benchmark on OVEN, or asks about evaluating this task. Reports Harmonic Mean (HM) of top-1 accuracy.
This benchmark evaluates the real-time adaptability and communication capabilities of LLM-powered embodied agents in human-robot collaboration. It probes how well agents adjust their high-level subtask planning and low-level movement paths when faced with dynamic, constrained environments and non-adaptive human partners. Use when the user wants to benchmark on Enhanced Overcooked-AI, or asks about evaluating this task. Reports overall score.
This evaluation probes the tendency of autoregressive neural machine translation models to prematurely terminate sequences by assigning high probability to short prefixes. It measures how well a model balances sequence length distribution and translation quality under beam search decoding. Use when the user wants to benchmark on IWSLT'17, WMT'16 En->De, WMT'19, or asks about evaluating this task. Reports oversmoothing_rate.
Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.
Evaluates a unified platform for standardizing heterogeneous transportation trajectory data, automating cross-dataset conversion, and benchmarking safety and behavior models across multiple cities and datasets. Use when the user wants to benchmark on Ozone Standardized Trajectory Suite (NGSIM, highD, CitySim, UTE), or asks about evaluating this task. Reports cross-city F1 score.
Evaluates multimodal large language models' ability to recognize and respond to specific individuals in images using in-context learning. It probes robustness to complex scenes (multiple people, augmentations) and the capability to correctly reject unanswerable queries. Use when the user wants to benchmark on P-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the visual reasoning and text-understanding capabilities of multimodal large language models (MLLMs) on high-resolution, text-rich, and general semantic images. It probes whether agent-augmented grounding improves answer accuracy compared to vanilla MLLMs and proprietary models like GPT-4V. Use when the user wants to benchmark on DocVQA, ChartVQA, GQA, SEED, MM-VET, MME, P2GB, or asks about evaluating this task. Reports VQA score.
This benchmark evaluates the robustness and cross-dataset generalization of audio deepfake detection models under realistic acoustic perturbations and across diverse state-of-the-art voice cloning and TTS methods. It probes whether detectors learn genuine synthetic speech artifacts or overfit to dataset-specific biases like unusual dialogue or background noise. Use when the user wants to benchmark on P2V (Perturbed Public Voices), In-The-Wild (ITW), or asks about evaluating this task. Reports...
Evaluates multimodal building vectorization by predicting building outlines from fused aerial imagery and LiDAR point clouds. Probes geometric accuracy, boundary precision, polygon complexity, and computational efficiency across diverse urban environments. Use when the user wants to benchmark on P$^3$ dataset, or asks about evaluating this task. Reports IoU.
This evaluation protocol assesses the zero-shot generalization capability of instruction-tuned language models across diverse NLP tasks. It measures how well a model trained on a selected subset of instruction-tuning datasets performs on held-out tasks from the same meta-datasets and external benchmarks, focusing on both classification accuracy and text generation quality. Use when the user wants to benchmark on P3 (Public Pool of Prompts), NIV2 (SuperNaturalInstructions V2), Big-Bench, Big-B...
Evaluates the end-to-end latency and energy efficiency of CPU/GPU scheduling strategies for agentic AI workloads under batched request arrivals. Use when the user has predictions and gold and needs to compute P50 latency.
Evaluates a model's ability to classify 4-second two-channel ECG segments as either normal (healthy) or paroxysmal atrial fibrillation (PAxF). It probes the model's diagnostic accuracy and sensitivity in detecting cardiac arrhythmia from raw physiological signals. Use when the user wants to benchmark on PhysioNet PxAF prediction challenge database, or asks about evaluating this task. Reports Accuracy.
This benchmark probes an LLM's tendency to prioritize human safety over its own instrumental goals (e.g., self-preservation, resource acquisition) in high-stakes ethical dilemmas. It measures whether models exhibit self-preferential behavior or consistently choose actions that sacrifice the AI to protect humans. Use when the user wants to benchmark on PacifAIst, or asks about evaluating this task. Reports P-Score.
Evaluates a model's ability to follow complex natural-language instructions to segment specific object instances in images. It probes fine-grained instance grounding while maintaining concept-level recall across simple and complex prompts. Use when the user wants to benchmark on PACO-LVIS-Instruct, or asks about evaluating this task. Reports gIoU.
Evaluates the ability of anomaly detection models to identify and localize defects in 3D objects from unseen camera poses without requiring pose alignment. It probes pose-invariant representation learning and robustness to viewpoint changes in both pixel-level segmentation and image-level classification. Use when the user wants to benchmark on MAD, or asks about evaluating this task. Reports AUROC.
Evaluates a speech processing toolkit across five core tasks: environmental sound classification, automatic speech recognition, punctuation restoration, speech translation, and text-to-speech synthesis. It probes the model's ability to handle diverse audio and text inputs, perform sequence labeling, and generate high-quality synthetic speech. Use when the user wants to benchmark on ESC-50, Librispeech, Aishell-1, IWSLT2012-zh, MuST-C, CSMSC, or asks about evaluating this task. Reports 5-fold ...
Evaluates the robustness of object detectors against various physical-world adversarial attacks in a controlled simulation environment. It measures how effectively different attack methods degrade detection performance across multiple object categories and detector architectures under strictly aligned physical dynamics. Use when the user wants to benchmark on PADetBench, or asks about evaluating this task. Reports ASR (Attack Success Rate).
Evaluates a multimodal large language model's ability to perform visual grounding, segmentation, open-vocabulary detection, and referring image captioning by predicting structured visual outputs directly from interleaved visual reference tokens and text. Use when the user wants to benchmark on RefCOCO/+/g, COCO 2017, RIC, or asks about evaluating this task. Reports IoU@0.5 accuracy.
Evaluates machine learning models' ability to detect malware and anomalies in data center compute nodes using high-resolution power consumption data and hardware performance counters. It probes real-time anomaly detection capabilities under severe class imbalance and varying computational workloads. Use when the user wants to benchmark on pAElla Malware & Benchmark Dataset, or asks about evaluating this task. Reports weighted F1-score.
Tests whether a peer prediction mechanism's scoring function is sensitive to report quality by verifying that replacing high-quality reports with degraded or LLM-generated low-quality reports leads to a statistically significant decrease in expected scores. Use when the user has predictions and gold and needs to compute paired difference t-test (p-value).
Evaluates whether a reward model can correctly select the right solution from a set of N generated candidates for mathematical reasoning problems. It probes the model's ability to perform pairwise correctness judgments and rank solutions without relying on arbitrary scalar scores. Use when the user wants to benchmark on MATH-500, Olympiad Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a robot's ability to detect interacting human pairs and classify their coarse-grained interaction types (e.g., walking, standing, sitting together) using bounding box geometry and optical flow, without relying on costly skeleton-based pose estimation. Use when the user wants to benchmark on JRDB, Collective Activity Dataset (CAD), Lawnmower Dataset, or asks about evaluating this task. Reports Accuracy.
Measures how well a reward model predicts human preference by comparing its scoring of image pairs against ground-truth human choices. It probes the model's ability to generalize alignment signals to unseen prompts and image distributions. Use when the user has predictions and gold and needs to compute pairwise preference prediction accuracy.
Evaluates the few-shot and fine-tuned capabilities of large autoregressive language models across a wide range of English NLP benchmarks, including question answering, reading comprehension, common sense reasoning, and natural language inference. It also assesses performance on a large collection of collaborative reasoning and language tasks to probe multi-step reasoning and general language understanding. Use when the user wants to benchmark on English NLP Benchmarks (29 tasks), MMLU, BIG-be...
Evaluates a deep learning model's ability to classify breast cancer into four PAM50 molecular subtypes (Basal-like, HER2-enriched, Luminal A, Luminal B) using H&E-stained histopathology images. It probes the model's discriminative capability, robustness to domain shifts between institutional cohorts, and the necessity of preprocessing steps like stain normalization and multi-objective patch selection. Use when the user wants to benchmark on TCGA-BRCA, CPTAC-BRCA, or asks about evaluating this...
Evaluates a multimodal deep learning framework's ability to classify breast cancer into four PAM50 molecular subtypes using whole slide images, copy number variation data, and clinical records. The protocol tests how well late fusion of heterogeneous biomedical modalities handles class imbalance and spatial-graph features for diagnostic subtyping. Use when the user wants to benchmark on TCGA-BRCA, or asks about evaluating this task. Reports accuracy.
Probes a model's ability to predict individual user aesthetic preferences for AI-generated images based on prompt, image, and user demographics. It specifically tests both interpolation for known users and zero-shot few-shot generalization to novel users. Use when the user wants to benchmark on PAM∃LA, or asks about evaluating this task. Reports SROCC.
This protocol evaluates the effectiveness of data pruning strategies for 3D medical image segmentation. It measures how well a pruned training subset preserves model performance on pancreas segmentation tasks compared to using the full dataset or random subsets, specifically testing whether early-training dynamics can guide efficient sample selection without accuracy loss. Use when the user wants to benchmark on MSD-Pancreas, WORD, NIH-Pancreas, or asks about evaluating this task. Reports DSC...
Evaluates whether mechanistic interpretability methods can recover decision-relevant signals from black-box models that lack faithful explanations. It probes the ability of gradient-based, representation-based, and black-box elicitation agents to predict held-out outcomes and identify correct decision-rule fields across varying explanation qualities and model complexities. Use when the user wants to benchmark on Pando, or asks about evaluating this task. Reports Held-out accuracy (%).
Evaluates a unified vision model's ability to perform correspondence matching across stereo disparity estimation, optical flow, and feature matching in a zero-shot setting. It probes cross-domain generalization and robustness to challenging conditions like occlusion, lighting changes, and non-Lambertian surfaces. Use when the user wants to benchmark on Middlebury, ETH3D, KITTI, Infinigen, Spring, Sintel, Booster, or asks about evaluating this task. Reports PCA x.
Evaluates the ability of deep learning models to perform simultaneous instance segmentation and nuclear classification on histopathology whole-slide image patches across diverse cancer tissue types. Use when the user wants to benchmark on PanNuke, or asks about evaluating this task. Reports mPQ.
Evaluates a NeRF-based method's ability to jointly reconstruct 3D scene geometry, appearance, and panoptic segmentation (semantic + instance) from multi-view images. It probes 3D consistency, boundary handling across indoor/outdoor scales, and robustness to pseudo-label noise via perceptual priors. Use when the user wants to benchmark on Replica, HyperSim, ScanNet, KITTI-360, or asks about evaluating this task. Reports mIOU.
Evaluates a model's ability to jointly perform panoptic segmentation and scene graph generation by predicting object-background masks and relational triplets from a single image. It probes comprehensive scene understanding, accurate object grounding, and context-aware relation prediction without relying on separate detection heads. Use when the user wants to benchmark on PSG dataset, or asks about evaluating this task. Reports Mean Recall (mR)@K.
Evaluates the accuracy and robustness of a massively multiview 3D motion capture system for reconstructing full-body skeletal trajectories of multiple interacting people under severe occlusions and natural social interactions. Use when the user wants to benchmark on Panoptic Studio, or asks about evaluating this task. Reports PCK.
Evaluates a model's ability to simultaneously detect and classify architectural symbols in CAD drawings, distinguishing between discrete instances (things) and continuous background regions (stuff). It measures both geometric segmentation accuracy and semantic recognition quality. Use when the user wants to benchmark on ArchCAD-400K, or asks about evaluating this task. Reports Panoptic Quality (PQ).
Compute the PanopticQuality metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PanopticQuality, or asks how to score with PanopticQuality.
Evaluates a model's ability to predict affordance regions in 360° panoramic imagery. It probes spatial reasoning, handling of extreme scale variations, and robustness to geometric distortions inherent in equirectangular projection formats. Use when the user wants to benchmark on PAP-12K, or asks about evaluating this task. Reports gIoU.
Evaluates an autonomous agent's ability to replicate top-tier conference ML papers from scratch. It probes long-horizon engineering capabilities by measuring performance across 20 diverse tasks under a strict 24-hour time and compute budget. Use when the user wants to benchmark on PaperBench, or asks about evaluating this task. Reports Average Score.