Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,185–5,208 of 23,574 skills
This benchmark evaluates deep learning models on skeleton-based human motion rehabilitation assessment. It probes both classification (e.g., motion type or health status) and extrinsic regression tasks using standardized cross-subject splits and min-max normalization to prevent data leakage. Use when the user wants to benchmark on Rehab-Pile, or asks about evaluating this task. Reports accuracy.
Evaluates online resource allocation algorithms by measuring their cumulative regret and cumulative unfairness across simulated environments with varying resource binding and degeneracy conditions. Use when the user has predictions and gold and needs to compute cumulative_unfairness.
Evaluates the ability of vision-language models to generate accurate and semantically rich chest X-ray radiology reports from medical images. It tests both lexical overlap and clinical semantic alignment of the generated findings and impressions against ground-truth reports. Use when the user wants to benchmark on ReXGradient-160K, or asks about evaluating this task. Reports COMET.
Evaluates whether increasing the length of a user's purchase history context improves recommendation quality for LLM-based agents. It probes the saturation point of personalization reasoning and the cost-efficiency trade-off of context length. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports quality scores.
Evaluates conversational recommender systems on next-item prediction and joint narrative generation, specifically testing how well models incorporate user interaction history and explicit natural language critiques to produce accurate recommendations and contextually grounded textual explanations. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports Recall@10.
Evaluates NLP models' ability to perform legal information extraction (NER) and predict refugee claim decision outcomes from Canadian legal documents. It probes the model's capacity to handle domain-specific terminology, extract structured entities from unstructured text, and classify case outcomes based on judicial reasoning. Use when the user wants to benchmark on Canadian Refugee Status Determination (RSD) Cases, or asks about evaluating this task. Reports accuracy.
Evaluates NLP models on summarizing student course reflections across three tasks: document selection, phrase extraction with support counts, and abstractive summarization. It probes specificity-aware summarization capabilities and model robustness in low-resource educational settings with variable text structure. Use when the user wants to benchmark on ReflectSumm, or asks about evaluating this task. Reports Unspecified in text.
Evaluates single-view 3D stereo reconstruction quality in scenes containing mirror reflections. It probes a model's ability to leverage virtual views generated from mirror reflections to recover accurate 3D geometry and camera poses, outperforming standard monocular or stereo baselines that typically hallucinate depth or collapse in reflective regions. Use when the user wants to benchmark on Synthetic Dataset (Reflect3r), Real-world Mirror Scenes, or asks about evaluating this task. Reports F...
Evaluates the zero-shot generalization capability of autoregressive language models across multiple task aggregates. It measures how well models trained on raw web data perform on downstream tasks without any fine-tuning or prompt engineering. Use when the user wants to benchmark on Eleuther AI LM evaluation harness (zero-shot aggregates), or asks about evaluating this task. Reports zero-shot accuracy.
This benchmark evaluates a model's ability to perform relation extraction on complex, domain-specific financial documents (SEC 10-X filings). It specifically probes challenges such as numerical inference, semantic ambiguity between similar relation types, and directional dependency resolution in long-range financial text. Use when the user wants to benchmark on REFiND, or asks about evaluating this task. Reports micro-F1.
Evaluates a model's ability to generate unambiguous, context-aware text descriptions for specific objects in an image, and to comprehend those descriptions by correctly localizing the target object via bounding box prediction. Use when the user wants to benchmark on G-Ref, UNC-Ref, or asks about evaluating this task. Reports precision@1.
Evaluates a model's ability to perform dense grounded understanding by localizing and segmenting specific objects in images and videos based on natural language instructions or referring expressions. Use when the user wants to benchmark on Ref-SAV, RefCOCO, RefCOCO+, RefCOCOg, MeVIS, Ref-YTVOS, ReVOS, or asks about evaluating this task. Reports cIoU.
Evaluates visual grounding capabilities by measuring how well a model segments or localizes objects in images based on natural language descriptions. It probes the model's ability to handle ambiguous references, diverse textual forms, and generalized referring expressions without task-specific decoders. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, gRefCOCO, or asks about evaluating this task. Reports mIoU.
This benchmark evaluates a model's ability to generate referring expressions that enable humans to quickly and accurately identify a target object in an image. It prioritizes human comprehension speed and accuracy over purely semantic correctness, particularly for low-salience targets. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, RefGTA, or asks about evaluating this task. Reports R1-CIDEr.
This evaluation probes a model's ability to understand visual context by localizing objects described by natural language expressions (comprehension) and generating unambiguous, context-aware descriptions for objects in an image (generation). It specifically tests whether the model leverages intra-category visual comparisons and joint language modeling to produce discriminative referring expressions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating...
Probes a model's ability to ground referring expressions in images by understanding spatial and relational context between objects. It evaluates weakly supervised region proposal scoring and context pooling mechanisms without requiring explicit bounding box annotations for context regions. Use when the user wants to benchmark on Google RefExp, UNC RefExp, or asks about evaluating this task. Reports Precision@1.
This benchmark evaluates a model's ability to perform pixel-level image segmentation conditioned on natural language expressions. It probes spatial reasoning, attribute grounding, and fine-grained visual-linguistic alignment by requiring the model to segment specific objects or amorphous regions described in text. Use when the user wants to benchmark on ReferIt, or asks about evaluating this task. Reports prec@0.5.
Evaluates Multimodal Large Language Models (MLLMs) on automatic sports refereeing tasks, probing their ability to detect incidents, classify fouls, apply sport-specific rules, and ground decisions temporally across 11 different sports. Use when the user wants to benchmark on RefereeBench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform referring expression segmentation at both object and part levels. It probes fine-grained cross-modal alignment and pixel-level semantic understanding by requiring precise mask prediction for diverse textual references. Use when the user wants to benchmark on RefCOCOm, or asks about evaluating this task. Reports mIoU.
Evaluates fine-grained visual grounding and spatial reasoning by measuring how accurately a multimodal model can localize regions in images corresponding to given referring expressions under varying linguistic and spatial conditions. Use when the user wants to benchmark on RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports IoU@50 accuracy.
This benchmark evaluates a model's ability to answer questions about chart images while simultaneously localizing the visual evidence (via bounding boxes) that supports the answer. It probes spatial-text alignment, arithmetic and logical reasoning over charts, and hallucination reduction through explicit grounding. Use when the user wants to benchmark on RefChartQA, or asks about evaluating this task. Reports answer accuracy.
Evaluates a model's ability to localize a target object in an aerial image based on a fine-grained natural language description. It probes cross-modal alignment, scale-invariant object detection, and handling of complex backgrounds with numerous distractors. Use when the user wants to benchmark on RefAerial, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports mP (average P@0.5/0.6/0.7/0.8).
This benchmark evaluates large language models' ability to detect, localize, and correct scientific confabulations in generated answers. It probes fine-grained factuality awareness, span-level error identification, and factual restoration capabilities under domain-specific scrutiny. Use when the user wants to benchmark on ReFACT, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to segment objects in audio-visual videos based on natural language referring expressions. It tests both seen categories and generalization to unseen categories, as well as handling null references where no object exists. Use when the user wants to benchmark on Ref-AVS Dataset, or asks about evaluating this task. Reports Jaccard Index ($\mathcal{J}$).