Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,953–5,976 of 23,907 skills
This benchmark evaluates an LLM's ability to accurately translate SQL queries across different database systems and dialects. It probes the model's capacity to handle system-specific syntax, semantic equivalence, and edge-case safeguards without relying on superficial string matching. Use when the user wants to benchmark on PARROT, or asks about evaluating this task. Reports Acc_EX.
This benchmark evaluates the ability of deep learning models to perform binary semantic segmentation of parking lots from satellite imagery. It specifically probes the model's capacity to generalize across different geographic locations and to leverage near-infrared (NIR) spectral data for improved contrast against vegetation and built environments. Use when the user wants to benchmark on ParkSeg12k, or asks about evaluating this task. Reports mIoU.
Evaluates autonomous cleaning robots' ability to navigate public park pathways, perceive and collect diverse litter types, avoid obstacles, and operate within strict physical and safety constraints. Use when the user wants to benchmark on Park Cleaning Benchmark, or asks about evaluating this task. Reports collected_items_weight_or_count.
Evaluates multilingual and multi-cultural LLM performance across 10 Indic languages using culturally nuanced prompts. It measures model quality via pairwise comparisons (Elo ratings) and direct assessment scores, while also analyzing human-LLM evaluator agreement and various biases (position, verbosity, self-bias). Use when the user wants to benchmark on PARIKSHA, or asks about evaluating this task. Reports Elo rating, Direct Assessment score.
This evaluation probes a model's ability to synthesize interpretable surrogate models (decision diagrams) that explicitly trade off prediction fidelity against structural simplicity. It measures how well a method can navigate the Pareto front between accuracy and explainability without collapsing them into a single weighted objective. Use when the user wants to benchmark on Airplane Perception Module (AP), Bank Loan Predictor (BL), Theorem Prover Solvability Predictor (TP), or asks about eval...
Evaluates how well speech-to-speech models adapt to and reflect paralinguistic styles (age, emotion, gender, sarcasm) in spoken responses. It measures both content appropriateness and stylistic alignment against ground-truth or human-annotated references. Use when the user wants to benchmark on ParaS2SBench, IEMOCAP, MELD, or asks about evaluating this task. Reports ParaS2SBench score.
Evaluates how different deep learning model architectures and hyperparameters affect hardware performance across TPU, GPU, and CPU platforms. It probes the interaction between model attributes (size, type, batch size) and hardware bottlenecks like memory bandwidth, compute utilization, and data infeed overhead. Use when the user wants to benchmark on ParaDnn, or asks about evaluating this task. Reports performance.
Evaluates a privacy-preserving LLM delegation pipeline that sanitizes user queries before sending them to a remote API model. It measures the trade-off between maintaining response quality and minimizing personally identifiable information (PII) leakage in the sanitized prompts. Use when the user wants to benchmark on PUPA-TNB, or asks about evaluating this task. Reports QUAL.
Evaluates how well a language model's generated responses align with a specific individual's personality traits across the Big Five and Dark Triad dimensions. It measures the distance between the model's predicted personality profile and the target profile, with lower scores indicating better alignment. Use when the user wants to benchmark on PAPI, or asks about evaluating this task. Reports Aligned Score.
Evaluates the accuracy and interaction efficiency of LLM agents performing multi-turn tool-use for scientific paper question-answering. It probes the model's ability to plan, execute tool calls, and extract answers from complex academic documents without excessive interaction or getting stuck in loops. Use when the user wants to benchmark on AirQA-Real, SciDQA, or asks about evaluating this task. Reports Avg., I-Avg.
Evaluates multimodal LLMs' ability to perform integrated agentic reasoning and critical assessment over scientific papers. It probes capabilities in multimodal grounding, experimental interpretation, cross-source evidence synthesis via tool use, and critical evaluation of research claims. Use when the user wants to benchmark on PaperMind, or asks about evaluating this task. Reports F1 score.
Evaluates an autonomous agent's ability to replicate top-tier conference ML papers from scratch. It probes long-horizon engineering capabilities by measuring performance across 20 diverse tasks under a strict 24-hour time and compute budget. Use when the user wants to benchmark on PaperBench, or asks about evaluating this task. Reports Average Score.
Evaluates a model's ability to predict affordance regions in 360° panoramic imagery. It probes spatial reasoning, handling of extreme scale variations, and robustness to geometric distortions inherent in equirectangular projection formats. Use when the user wants to benchmark on PAP-12K, or asks about evaluating this task. Reports gIoU.
Evaluates a model's ability to simultaneously detect and classify architectural symbols in CAD drawings, distinguishing between discrete instances (things) and continuous background regions (stuff). It measures both geometric segmentation accuracy and semantic recognition quality. Use when the user wants to benchmark on ArchCAD-400K, or asks about evaluating this task. Reports Panoptic Quality (PQ).
Evaluates the accuracy and robustness of a massively multiview 3D motion capture system for reconstructing full-body skeletal trajectories of multiple interacting people under severe occlusions and natural social interactions. Use when the user wants to benchmark on Panoptic Studio, or asks about evaluating this task. Reports PCK.
Evaluates a model's ability to jointly perform panoptic segmentation and scene graph generation by predicting object-background masks and relational triplets from a single image. It probes comprehensive scene understanding, accurate object grounding, and context-aware relation prediction without relying on separate detection heads. Use when the user wants to benchmark on PSG dataset, or asks about evaluating this task. Reports Mean Recall (mR)@K.
Evaluates a NeRF-based method's ability to jointly reconstruct 3D scene geometry, appearance, and panoptic segmentation (semantic + instance) from multi-view images. It probes 3D consistency, boundary handling across indoor/outdoor scales, and robustness to pseudo-label noise via perceptual priors. Use when the user wants to benchmark on Replica, HyperSim, ScanNet, KITTI-360, or asks about evaluating this task. Reports mIOU.
Evaluates the ability of deep learning models to perform simultaneous instance segmentation and nuclear classification on histopathology whole-slide image patches across diverse cancer tissue types. Use when the user wants to benchmark on PanNuke, or asks about evaluating this task. Reports mPQ.
Evaluates a unified vision model's ability to perform correspondence matching across stereo disparity estimation, optical flow, and feature matching in a zero-shot setting. It probes cross-domain generalization and robustness to challenging conditions like occlusion, lighting changes, and non-Lambertian surfaces. Use when the user wants to benchmark on Middlebury, ETH3D, KITTI, Infinigen, Spring, Sintel, Booster, or asks about evaluating this task. Reports PCA x.
Evaluates whether mechanistic interpretability methods can recover decision-relevant signals from black-box models that lack faithful explanations. It probes the ability of gradient-based, representation-based, and black-box elicitation agents to predict held-out outcomes and identify correct decision-rule fields across varying explanation qualities and model complexities. Use when the user wants to benchmark on Pando, or asks about evaluating this task. Reports Held-out accuracy (%).
This protocol evaluates the effectiveness of data pruning strategies for 3D medical image segmentation. It measures how well a pruned training subset preserves model performance on pancreas segmentation tasks compared to using the full dataset or random subsets, specifically testing whether early-training dynamics can guide efficient sample selection without accuracy loss. Use when the user wants to benchmark on MSD-Pancreas, WORD, NIH-Pancreas, or asks about evaluating this task. Reports DSC...
Probes a model's ability to predict individual user aesthetic preferences for AI-generated images based on prompt, image, and user demographics. It specifically tests both interpolation for known users and zero-shot few-shot generalization to novel users. Use when the user wants to benchmark on PAM∃LA, or asks about evaluating this task. Reports SROCC.
Evaluates a multimodal deep learning framework's ability to classify breast cancer into four PAM50 molecular subtypes using whole slide images, copy number variation data, and clinical records. The protocol tests how well late fusion of heterogeneous biomedical modalities handles class imbalance and spatial-graph features for diagnostic subtyping. Use when the user wants to benchmark on TCGA-BRCA, or asks about evaluating this task. Reports accuracy.
Evaluates a deep learning model's ability to classify breast cancer into four PAM50 molecular subtypes (Basal-like, HER2-enriched, Luminal A, Luminal B) using H&E-stained histopathology images. It probes the model's discriminative capability, robustness to domain shifts between institutional cohorts, and the necessity of preprocessing steps like stain normalization and multi-objective patch selection. Use when the user wants to benchmark on TCGA-BRCA, CPTAC-BRCA, or asks about evaluating this...