Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,945–7,968 of 20,853 skills
Evaluates dense prompt following on multi-requirement prompts by decomposing them into dependency-structured VQA checks spanning entity presence, attributes, relations, and counts. Use when the user wants to benchmark on DPG-Bench, or asks about evaluating this task. Reports Overall.
Evaluates optical flow estimation models on standard and high-resolution benchmarks to measure accuracy, generalization across input sizes, and robustness to resolution scaling without tiling. It tests how well adaptive pyramid architectures maintain prediction stability when input dimensions increase from 1K to 8K. Use when the user wants to benchmark on Spring, Middlebury-ST, VIPER, Kubric-NK, MPI-Sintel, KITTI 2015, or asks about evaluating this task. Reports EPE.
This evaluation protocol assesses the utility and privacy-utility trade-off of Graph Neural Networks trained with Differentially Private Stochastic Gradient Descent (DP-SGD) on graph-level classification tasks. It probes whether formal privacy guarantees can be maintained across diverse graph structures (molecules, fingerprints, ECG signals, synthetic graphs) without severely degrading predictive performance compared to non-private baselines. Use when the user wants to benchmark on Synthetic,...
This evaluation probes how reliably language model scaling laws predict performance in over-trained regimes, where models are trained with significantly more tokens than parameters. It measures both next-token prediction accuracy on a held-out corpus and generalization across a broad suite of downstream zero-shot and few-shot tasks. Use when the user wants to benchmark on C4 eval, LLM-foundry, or asks about evaluating this task. Reports Validation loss.
Evaluates the downstream language and symbolic capabilities of LLMs pretrained on filtered web corpora. It probes general knowledge, reasoning, comprehension, and code/math problem-solving to assess how different data filtering strategies impact model performance. Use when the user wants to benchmark on Dolma (v1.6), Pile-github, or asks about evaluating this task. Reports average normalized accuracy.
This benchmark evaluates instruction-following capabilities of speech-large language models (SLLMs) by comparing performance when prompted with text versus spoken audio across nine diverse tasks. It probes cross-lingual generalization, prompt style robustness, and the model's ability to handle both text and speech modalities for input and output. Use when the user wants to benchmark on DOWIS (Do What I Say), FLEURS, MCIF, YTSeg, or asks about evaluating this task. Reports WER.
This evaluation probes the robustness and prompt sensitivity of large language models on multiple-choice benchmarks by measuring how performance varies across hundreds of millions of intent-preserving prompt perturbations across multiple dimensions. Use when the user wants to benchmark on DOVE, or asks about evaluating this task. Reports Accuracy.
Evaluates cross-domain generalization of mammography classification models under domain shift, specifically testing resilience to variations in pixel intensity distributions across different imaging devices and datasets. Use when the user wants to benchmark on NYU, HCTP, VinDr, CSAW, or asks about evaluating this task. Reports PR-AUC.
Evaluates multimodal large language models' ability to understand and reason about object orientation across four dimensions: frontal alignment, rotational transformations, relative directional relationships, and canonical orientation. It distinguishes between coarse categorical judgments and fine-grained angular estimations to probe 3D spatial reasoning capabilities. Use when the user wants to benchmark on DORI, or asks about evaluating this task. Reports accuracy.
Evaluates RAG-based question answering systems on defense-domain documents, measuring both retrieval effectiveness and end-to-end QA performance including task success, faithfulness, and generation quality. Use when the user wants to benchmark on DoRA, or asks about evaluating this task. Reports task-success.
Evaluates a robot policy's ability to perform manipulation tasks in environments with moving objects and dynamic spatiotemporal changes. It probes the model's capacity for historical context integration and future state anticipation to maintain control stability and task success under motion. Use when the user wants to benchmark on DOMINO@0.1, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates multi-source domain adaptation methods on image classification tasks across multiple domains with varying visual styles and categories. Use when the user wants to benchmark on DomainNet, or asks about evaluating this task. Reports average accuracy.
Evaluates the effectiveness and efficiency of software model checking configurations that use domain types to select abstract domains (BDD vs explicit-value) for variable abstraction. It probes how well different abstraction strategies handle verification tasks across various benchmark suites. Use when the user wants to benchmark on SV-COMP and RERS benchmark sets (SYSTEMC, ECA, LOCK, PRODUCT SIMULATOR, NTDRIVERS, SSH), or asks about evaluating this task. Reports Effectiveness.
This benchmark evaluates a model's out-of-distribution (OOD) generalization capability across multiple domain-shift datasets. It specifically probes whether models rely on true domain-invariant features learned from training domains versus leaking test-domain information through ImageNet pretraining weights or oracle hyperparameter selection. The protocol mandates training from scratch without pretrained weights and evaluating across multiple test domains to ensure a fair comparison of OOD ge...
Evaluates robotic policies on dynamic object manipulation, measuring their ability to react to moving objects, perceive visual/spatial/motion cues, and generalize across novel objects, scenes, and motion patterns. It specifically probes closed-loop reactivity, dynamic adaptation, long-horizon sequencing, and robustness to disturbances. Use when the user wants to benchmark on DOM, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates the natural language generation capabilities of models across 13 diverse Arabic tasks, including machine translation, summarization, question generation, and dialectal normalization. It probes how well models handle linguistic variability across Classical Arabic, Modern Standard Arabic, dialects, and Arabizi, as well as cross-lingual and code-switched scenarios. Use when the user wants to benchmark on Dolphin, or asks about evaluating this task. Reports BLEU.
Evaluates a model's capability to perform document understanding tasks, including key information extraction, visual question answering, table question answering, and structural reading comprehension, using textual layout representations derived from OCR and spatial verbalization. Use when the user wants to benchmark on DocVQA, InfographicsVQA, WikiTableQuestions, TabFact, SROIE, or asks about evaluating this task. Reports ANLS, accuracy.
Evaluates end-to-end document parsing models on their ability to extract structured content (text, formulas, tables, reading order) from both standardized printed documents and real-world captured images. It measures structural fidelity, multilingual robustness, and decoding stability under visual degradation. Use when the user wants to benchmark on OmniDocBench, XFUND, Wild-OmniDocBench, or asks about evaluating this task. Reports Overall.
Evaluates the ability of models to extract key information fields from visually-rich documents (receipts and forms). It probes generalization to unseen document templates by comparing performance on original versus resampled test splits. Use when the user wants to benchmark on SROIE, FUNSD, or asks about evaluating this task. Reports F1.
This benchmark evaluates Vision Language Models' ability to retrieve specific textual or multimodal "needles" embedded within long documents ranging from 5 to 200 pages. It probes visual-text alignment, long-context retrieval capabilities, and performance degradation as document length and token consumption increase. Use when the user wants to benchmark on Document Haystack, or asks about evaluating this task. Reports Accuracy.
This evaluation probes a model's ability to recommend specialist doctors for patients using implicit interaction data and limited demographic metadata. It specifically tests performance in both warm-start (seen patients) and cold-start (new patients) scenarios, emphasizing the model's capacity to handle popularity bias and recommend less popular specialists. Use when the user wants to benchmark on Doctor Recommendation Dataset, or asks about evaluating this task. Reports PS-nDCG@3.
Evaluates document-level relation extraction systems on multi-sentence reasoning, entity coreference resolution, and long-range dependency modeling. It measures how well models can predict relational facts between entity pairs across entire documents, including cases requiring evidence from multiple sentences. Use when the user wants to benchmark on DocRED, or asks about evaluating this task. Reports F1.
Evaluates the ability of object detection models to accurately identify and localize 11 distinct document layout elements (e.g., text, tables, figures, headers) on scanned or digital document pages. It measures robustness across diverse, real-world document types and tests how data splitting strategies and label definitions impact prediction accuracy. Use when the user wants to benchmark on DocLayNet, or asks about evaluating this task. Reports mAP@0.5-0.95.
Evaluates the performance overhead of Docker containers on deep learning workloads by benchmarking CPU, GPU, I/O, and training speed of representative neural networks (FCN, CNN, RNN) across different frameworks. Use when the user wants to benchmark on MNIST, Cifar10, PTB, or asks about evaluating this task. Reports second per batch.