Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 8,809–8,832 of 21,231 skills
Evaluates continual instruction tuning in multimodal large language models by measuring how well they retain task-specific instruction alignment and underlying reasoning knowledge when trained sequentially on diverse datasets. Use when the user wants to benchmark on ScienceQA, TextVQA, ImageNet, GQA, VizWiz, Grounding, VQAv2, OCR-VQA, or asks about evaluating this task. Reports Truth Alignment.
Evaluates an agent's ability to navigate to a specific target instance in multi-instance scenes through collaborative, open-ended dialogues with a human or simulated user. It probes the agent's uncertainty-aware reasoning, dialogue efficiency, and generalization to unseen object categories. Use when the user wants to benchmark on CoIN-Bench, IDKVQA, or asks about evaluating this task. Reports SR.
Evaluates large language models and fine-tuned baselines on multi-label medical text classification for clinical cohort recruitment. It probes the model's ability to extract and predict disease labels from unstructured radiology reports using few-shot prompting and knowledge graph augmentation. Use when the user wants to benchmark on IU-RR, MIMIC-CXR, or asks about evaluating this task. Reports F1-Score (F).
Evaluates whether neural coherence models can distinguish coherent text from artificially incoherent permutations and whether their scores correlate with human judgments on real-world downstream tasks like machine translation and summarization. Use when the user wants to benchmark on WSJ, WMT2017-2018, CNN/DM, DUC 2003, or asks about evaluating this task. Reports accuracy.
Probes compositional generalization in semantic parsing by evaluating whether models can correctly map out-of-distribution natural language sentences to their corresponding lambda calculus semantic representations. It specifically tests structural generalizations like argument role reversal, depth generalization, and voice transformation, as well as lexical generalizations. Use when the user wants to benchmark on COGS, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a neural machine translation model's ability to generalize compositionally by translating novel compound phrases that were not seen during training. It probes whether the model can correctly assemble semantic components in new syntactic contexts, revealing gaps between standard sentence-level fluency metrics and actual compositional robustness. Use when the user wants to benchmark on CoGnition, or asks about evaluating this task. Reports compound translation error rate.
Evaluates vision-language models on multi-page document understanding, specifically testing long-context compositional reasoning, fine-grained information extraction from forms, complex layout and chart comprehension, and cross-page navigation for answer localization. Use when the user wants to benchmark on MMLongbench-Doc, DUDE, SlideVQA, MP-DocVQA, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of retrieval-augmented language models to mitigate knowledge hallucination and maintain robustness in reading comprehension and question-answering tasks when processing long, noisy contexts with selective highlighting. Use when the user wants to benchmark on FELM, RACE-H, RACE-M, Natural Questions, TriviaQA, WebQ, or asks about evaluating this task. Reports F1 score.
Evaluates a model's ability to follow task-specific and composite text editing instructions across diverse tasks like grammar correction, simplification, coherence, style transfer, and paraphrasing. It also assesses generalization to unseen instructions and human-perceived writing efficiency. Use when the user wants to benchmark on JFLEG, TurkCorpus, ASSET, ITER (Coherence split), DISCOFUSE, ITER (Iterative text revision), GYAFC, WNC, MRPC, STS, QQP, or asks about evaluating this task. Report...
This benchmark evaluates a model's ability to make context-aware turn-taking decisions in multi-turn dialogues. It probes whether models can correctly predict one of four functional actions based on dialogue history and the current system state. The evaluation further diagnoses performance across 14 fine-grained interactional scenarios to reveal semantic misalignments beyond binary end-of-utterance detection. Use when the user wants to benchmark on CoDeTT, or asks about evaluating this task. ...
This benchmark evaluates an agent's ability to localize the onset of failure within long-horizon code execution trajectories by analyzing heterogeneous run artifacts. It probes how well models can distinguish genuinely failure-relevant steps from salient but irrelevant logs, diagnose execution bottlenecks, and recover from early wrong commitments under constrained token budgets. Use when the user wants to benchmark on CodeTraceBench, or asks about evaluating this task. Reports step-level F1.
This benchmark probes the robustness of machine translation systems to dialectal variations by measuring how consistently they translate semantically similar sentences in standard vs. dialectal forms. It evaluates whether models maintain translation quality and coherence when exposed to lexical and morphosyntactic variations across multiple languages. Use when the user wants to benchmark on CODET, or asks about evaluating this task. Reports COMET.
This benchmark probes large language models' ability to comprehend implicit code review intent by decomposing the task into change type recognition, change localization, and solution identification. It uses multiple-choice questions to evaluate whether models can accurately interpret pre-change code and reviewer comments without relying on surface-level code generation. Use when the user wants to benchmark on CodeReviewQA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify source code into the programming problem it was submitted to solve. It probes code representation learning and structural understanding by mapping code samples to their corresponding problem classes. Use when the user wants to benchmark on CodeNet (Java250, Python800, C++1000, C++1400), or asks about evaluating this task. Reports accuracy.
Evaluates code understanding and reasoning capabilities of large language models using a multiple-choice question format. It probes syntactic knowledge, semantic comprehension, and real-world software engineering problem-solving, revealing gaps in true code reasoning compared to open-ended generation benchmarks. Use when the user wants to benchmark on CodeMMLU, or asks about evaluating this task. Reports accuracy %.
Evaluates large language models' ability to process and generate code-mixed text across 18 languages and 8 distinct tasks. It probes cross-lingual reasoning, traditional NLP capabilities, and few-shot learning robustness when linguistic families are mixed within a single prompt. Use when the user wants to benchmark on CodeMixBench, or asks about evaluating this task. Reports accuracy.
This evaluation probes an LLM's ability to automatically detect and fix bugs in C programs by generating correct patches. It measures how effectively the model leverages test feedback, fault localization scores, and iterative reasoning to pass all provided test cases for each buggy submission. Use when the user wants to benchmark on Codeflaws, or asks about evaluating this task. Reports Repair Accuracy.
Evaluates the validity of the CodeBLEU metric for code synthesis by measuring its correlation with human programmer judgments across text-to-code generation, code translation, and code refinement tasks. Use when the user has predictions and gold and needs to compute CodeBLEU.
Evaluates code LLMs' alignment with human preferences on real-world, non-algorithmic coding tasks. It measures how well model-generated code matches human-like quality and preference compared to a baseline, rather than just syntactic or execution correctness. Use when the user wants to benchmark on CodeArena, EvalPlus, MultiPL-E, or asks about evaluating this task. Reports Pass@1, Win rate.
Evaluates a model's ability to perform test-time adaptation (TTA) for 3D object detection and autonomous driving tasks under domain shifts and sensor corruptions. It probes robustness to environmental changes (weather, lighting) and hardware failures without full retraining. Use when the user wants to benchmark on KITTI, KITTI-C, Waymo, nuScenes, nuScenes-C, or asks about evaluating this task. Reports NDS.
Evaluates Large Vision-Language Models on real-world autonomous driving corner cases across three tasks: general perception, regional perception, and driving suggestions. It probes the model's ability to accurately identify traffic-relevant objects, explain their impact on driving behavior, and generate actionable, rational driving advice in complex scenarios. Use when the user wants to benchmark on CODA-LM, or asks about evaluating this task. Reports Text-Score.
Evaluates a training-free, constraint-based data augmentation framework for low-resource NLP. It probes whether synthetically augmented data improves downstream performance across sequence classification, intent classification, named entity recognition, and question answering tasks compared to gold-only and other augmentation baselines. Use when the user wants to benchmark on Huffpost, Yahoo, OTS, ATIS, Massive, ConLL-2003, OntoNotes-5.0, EBMNLP, BC2GM, SQuAD, NewsQA, or asks about evaluating...
Evaluates object detection performance of scalable neural backbones on resource-constrained edge devices by measuring mean Average Precision across varying computational budgets and low input resolutions. Use when the user wants to benchmark on MS COCO, VOC2012, or asks about evaluating this task. Reports mAP.
Evaluates a model's ability to perform universal few-shot instance perception across object detection, instance segmentation, pose estimation, and object counting. It probes task-agnostic generalization and robustness in extremely low-shot (1-shot and 5-shot) scenarios, including unseen-task generalization for counting. Use when the user wants to benchmark on COCO-UniFS, PASCAL-5i, or asks about evaluating this task. Reports Det. AP.