Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,977–6,000 of 23,912 skills
Evaluates whether mechanistic interpretability methods can recover decision-relevant signals from black-box models that lack faithful explanations. It probes the ability of gradient-based, representation-based, and black-box elicitation agents to predict held-out outcomes and identify correct decision-rule fields across varying explanation qualities and model complexities. Use when the user wants to benchmark on Pando, or asks about evaluating this task. Reports Held-out accuracy (%).
This protocol evaluates the effectiveness of data pruning strategies for 3D medical image segmentation. It measures how well a pruned training subset preserves model performance on pancreas segmentation tasks compared to using the full dataset or random subsets, specifically testing whether early-training dynamics can guide efficient sample selection without accuracy loss. Use when the user wants to benchmark on MSD-Pancreas, WORD, NIH-Pancreas, or asks about evaluating this task. Reports DSC...
Probes a model's ability to predict individual user aesthetic preferences for AI-generated images based on prompt, image, and user demographics. It specifically tests both interpolation for known users and zero-shot few-shot generalization to novel users. Use when the user wants to benchmark on PAM∃LA, or asks about evaluating this task. Reports SROCC.
Evaluates a multimodal deep learning framework's ability to classify breast cancer into four PAM50 molecular subtypes using whole slide images, copy number variation data, and clinical records. The protocol tests how well late fusion of heterogeneous biomedical modalities handles class imbalance and spatial-graph features for diagnostic subtyping. Use when the user wants to benchmark on TCGA-BRCA, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to classify breast cancer into four PAM50 molecular subtypes (Basal-like, HER2-enriched, Luminal A, Luminal B) using H&E-stained histopathology images. It probes the model's discriminative capability, robustness to domain shifts between institutional cohorts, and the necessity of preprocessing steps like stain normalization and multi-objective patch selection. Use when the user wants to benchmark on TCGA-BRCA, CPTAC-BRCA, or asks about evaluating this...
Evaluates the few-shot and fine-tuned capabilities of large autoregressive language models across a wide range of English NLP benchmarks, including question answering, reading comprehension, common sense reasoning, and natural language inference. It also assesses performance on a large collection of collaborative reasoning and language tasks to probe multi-step reasoning and general language understanding. Use when the user wants to benchmark on English NLP Benchmarks (29 tasks), MMLU, BIG-be...
Measures how well a reward model predicts human preference by comparing its scoring of image pairs against ground-truth human choices. It probes the model's ability to generalize alignment signals to unseen prompts and image distributions. Use when the user has predictions and gold and needs to compute pairwise preference prediction accuracy.
Evaluates a robot's ability to detect interacting human pairs and classify their coarse-grained interaction types (e.g., walking, standing, sitting together) using bounding box geometry and optical flow, without relying on costly skeleton-based pose estimation. Use when the user wants to benchmark on JRDB, Collective Activity Dataset (CAD), Lawnmower Dataset, or asks about evaluating this task. Reports Accuracy.
Evaluates whether a reward model can correctly select the right solution from a set of N generated candidates for mathematical reasoning problems. It probes the model's ability to perform pairwise correctness judgments and rank solutions without relying on arbitrary scalar scores. Use when the user wants to benchmark on MATH-500, Olympiad Bench, or asks about evaluating this task. Reports accuracy.
Tests whether a peer prediction mechanism's scoring function is sensitive to report quality by verifying that replacing high-quality reports with degraded or LLM-generated low-quality reports leads to a statistically significant decrease in expected scores. Use when the user has predictions and gold and needs to compute paired difference t-test (p-value).
Evaluates a multimodal large language model's ability to perform visual grounding, segmentation, open-vocabulary detection, and referring image captioning by predicting structured visual outputs directly from interleaved visual reference tokens and text. Use when the user wants to benchmark on RefCOCO/+/g, COCO 2017, RIC, or asks about evaluating this task. Reports IoU@0.5 accuracy.
Evaluates the robustness of object detectors against various physical-world adversarial attacks in a controlled simulation environment. It measures how effectively different attack methods degrade detection performance across multiple object categories and detector architectures under strictly aligned physical dynamics. Use when the user wants to benchmark on PADetBench, or asks about evaluating this task. Reports ASR (Attack Success Rate).
Evaluates a speech processing toolkit across five core tasks: environmental sound classification, automatic speech recognition, punctuation restoration, speech translation, and text-to-speech synthesis. It probes the model's ability to handle diverse audio and text inputs, perform sequence labeling, and generate high-quality synthetic speech. Use when the user wants to benchmark on ESC-50, Librispeech, Aishell-1, IWSLT2012-zh, MuST-C, CSMSC, or asks about evaluating this task. Reports 5-fold ...
Evaluates the ability of anomaly detection models to identify and localize defects in 3D objects from unseen camera poses without requiring pose alignment. It probes pose-invariant representation learning and robustness to viewpoint changes in both pixel-level segmentation and image-level classification. Use when the user wants to benchmark on MAD, or asks about evaluating this task. Reports AUROC.
Evaluates a model's ability to follow complex natural-language instructions to segment specific object instances in images. It probes fine-grained instance grounding while maintaining concept-level recall across simple and complex prompts. Use when the user wants to benchmark on PACO-LVIS-Instruct, or asks about evaluating this task. Reports gIoU.
This benchmark probes an LLM's tendency to prioritize human safety over its own instrumental goals (e.g., self-preservation, resource acquisition) in high-stakes ethical dilemmas. It measures whether models exhibit self-preferential behavior or consistently choose actions that sacrifice the AI to protect humans. Use when the user wants to benchmark on PacifAIst, or asks about evaluating this task. Reports P-Score.
Evaluates a model's ability to classify 4-second two-channel ECG segments as either normal (healthy) or paroxysmal atrial fibrillation (PAxF). It probes the model's diagnostic accuracy and sensitivity in detecting cardiac arrhythmia from raw physiological signals. Use when the user wants to benchmark on PhysioNet PxAF prediction challenge database, or asks about evaluating this task. Reports Accuracy.
This evaluation protocol assesses the zero-shot generalization capability of instruction-tuned language models across diverse NLP tasks. It measures how well a model trained on a selected subset of instruction-tuning datasets performs on held-out tasks from the same meta-datasets and external benchmarks, focusing on both classification accuracy and text generation quality. Use when the user wants to benchmark on P3 (Public Pool of Prompts), NIV2 (SuperNaturalInstructions V2), Big-Bench, Big-B...
Evaluates multimodal building vectorization by predicting building outlines from fused aerial imagery and LiDAR point clouds. Probes geometric accuracy, boundary precision, polygon complexity, and computational efficiency across diverse urban environments. Use when the user wants to benchmark on P$^3$ dataset, or asks about evaluating this task. Reports IoU.
This benchmark evaluates the robustness and cross-dataset generalization of audio deepfake detection models under realistic acoustic perturbations and across diverse state-of-the-art voice cloning and TTS methods. It probes whether detectors learn genuine synthetic speech artifacts or overfit to dataset-specific biases like unusual dialogue or background noise. Use when the user wants to benchmark on P2V (Perturbed Public Voices), In-The-Wild (ITW), or asks about evaluating this task. Reports...
Evaluates the visual reasoning and text-understanding capabilities of multimodal large language models (MLLMs) on high-resolution, text-rich, and general semantic images. It probes whether agent-augmented grounding improves answer accuracy compared to vanilla MLLMs and proprietary models like GPT-4V. Use when the user wants to benchmark on DocVQA, ChartVQA, GQA, SEED, MM-VET, MME, P2GB, or asks about evaluating this task. Reports VQA score.
Evaluates multimodal large language models' ability to recognize and respond to specific individuals in images using in-context learning. It probes robustness to complex scenes (multiple people, augmentations) and the capability to correctly reject unanswerable queries. Use when the user wants to benchmark on P-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a unified platform for standardizing heterogeneous transportation trajectory data, automating cross-dataset conversion, and benchmarking safety and behavior models across multiple cities and datasets. Use when the user wants to benchmark on Ozone Standardized Trajectory Suite (NGSIM, highD, CitySim, UTE), or asks about evaluating this task. Reports cross-city F1 score.
Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.