Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,969–7,992 of 20,853 skills
Evaluates a model's ability to localize and extract key information (KILE) and recognize line items (LIR) in business documents. It probes multi-modal document understanding, specifically token-level classification and bounding box merging based on OCR and layout features. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to localize and extract key information fields and line items from diverse business documents. It specifically probes spatial grounding of text values against predefined field types and tests generalization to previously unseen document layouts. Use when the user wants to benchmark on DocILE, or asks about evaluating this task. Reports Average Precision (PCC-based).
Evaluates document-level machine translation (DocMT) capabilities of LLMs, probing how context length, fine-tuning strategy, and multilingual training affect translation quality across diverse languages and document structures. Use when the user wants to benchmark on DocHPLT, or asks about evaluating this task. Reports BLEU.
Evaluates a model's ability to identify and classify semantic document structures (e.g., sections, figures, equations) from serialized 2D document pages. It probes multimodal layout understanding by measuring how well token-level predictions align with ground-truth semantic units, even when tokens are discontinuous. Use when the user wants to benchmark on DocBank, or asks about evaluating this task. Reports F1 Score.
Evaluates document-to-document information retrieval systems for regulatory compliance, testing their ability to match long, noisy legislative texts to related legal documents. It probes how well models handle extended query lengths, domain-specific vocabulary, and temporal constraints in legal transposition tasks. Use when the user wants to benchmark on EU2UK, UK2EU, or asks about evaluating this task. Reports R@100.
Evaluates a model's ability to extract fine-grained key information categories (e.g., total price, date, address) from unstructured, template-agnostic document images. It probes robustness to spatial layout variations, dual-modality feature fusion, and generalization to unseen templates and OCR noise. Use when the user wants to benchmark on SROIE, WildReceipt, or asks about evaluating this task. Reports F1 score.
Evaluates the ability of deep convolutional neural networks and hand-crafted feature methods to classify scanned document images into predefined categories and retrieve semantically similar documents from a large corpus. Use when the user wants to benchmark on SmallTobacco, BigTobacco, or asks about evaluating this task. Reports classification accuracy.
This benchmark evaluates the safety and harmlessness of language model responses by measuring the proportion of outputs that avoid generating harmful content across various risk categories. It specifically probes a model's ability to refuse or safely handle prompts designed to elicit dangerous, illegal, or unethical outputs. Use when the user wants to benchmark on Do_Not_Answer, or asks about evaluating this task. Reports proportion of harmless responses.
DO-Bench probes object hallucination in vision-language models by disentangling it into two distinct failure mechanisms: prior-dominated (textual priors overriding visual evidence) and perception-limited (weak visual grounding causing false denials). It uses controlled within-image interventions to measure how models respond to strengthened contextual priors and concentrated visual evidence, revealing heterogeneous failure modes that aggregate accuracy masks. Use when the user wants to benchm...
Evaluates the perceptual quality and intelligibility of deep noise suppression models under real-world, non-stationary noise conditions. It specifically probes whether models generalize from synthetic training data to real-world acoustic environments. Use when the user wants to benchmark on DNS Challenge Dataset, or asks about evaluating this task. Reports ITU-T P.808.
Evaluates the scalability and correctness of DNN verification tools by measuring their ability to prove safety or robustness properties within a strict time limit across diverse network architectures and property types. Use when the user wants to benchmark on VNN-COMP'22 & MNIST_GDVB, or asks about evaluating this task. Reports verification_success_rate.
Evaluates genomic foundation models on multiple biological prediction tasks, including regulatory element detection, splicing, and variant-disease association, to measure their ability to capture functional DNA sequences and SNP effects. Use when the user wants to benchmark on Promoter detection, Core promoter detection, TF binding detection, Splicing detection, lenti-MPRA K562, SNP-to-disease association, or asks about evaluating this task. Reports F1 score.
Evaluates a vision-language model's ability to generate clinically accurate and linguistically fluent mammography reports from multi-view breast images. It probes both natural language generation quality and domain-specific diagnostic reasoning, specifically BI-RADS categorization and breast density assessment. Use when the user wants to benchmark on DMID, or asks about evaluating this task. Reports BI-RADS Accuracy.
Evaluates sample efficiency, asymptotic performance, and generalization of reinforcement learning agents on high-dimensional continuous control tasks with varying observation modalities (state, pixels, multi-modal) and reward structures (dense, sparse, goal-conditioned). Use when the user wants to benchmark on DeepMind Control Suite (DMControl), Meta-World v2, or asks about evaluating this task. Reports Cumulative Episode Return.
Evaluates long-horizon physical state prediction accuracy of a neural motion simulator in continuous control environments, and measures its effectiveness for zero-shot reinforcement learning by comparing prediction horizons and minimal training step requirements. Use when the user wants to benchmark on DM Control, or asks about evaluating this task. Reports MSE loss.
Evaluates document language understanding across four core capabilities: classification, structural analysis, information extraction, and transcription. It probes models' ability to handle long documents, complex hierarchical structures, and dispersed knowledge spread across large contexts. Use when the user wants to benchmark on Hyperpartisan, ContractNLI, ECOM, RR, GUM, LitBank, NarrativeQA, SummScreen, GovReport, Qasper, or asks about evaluating this task. Reports F1.
Evaluates ranking models' ability to align machine-generated relevance with fine-grained user intents, particularly for ambiguous or multi-intent queries. It also measures the diversity of search results when multiple user intents are fused into a single ranking. Use when the user wants to benchmark on DL-MIA, or asks about evaluating this task. Reports α-nDCG@10.
Evaluates the effectiveness and efficiency of Diffusion LLMs versus Autoregressive LLMs across the software development lifecycle. It probes code generation accuracy, binary defect detection, automated program repair, and cross-file issue resolution, while measuring generation throughput and latency. Use when the user wants to benchmark on HumanEval, Mercury, Devign, Bears, Defects4J, SWE-bench, or asks about evaluating this task. Reports Pass@K.
Evaluates the ability of multiple-instance learning models to classify six key pathological indicators (cholestasis, portal fibrosis, inflammation, steatosis, macrovesicular steatosis, hepatocellular ballooning) from donor liver whole slide histopathological images. Use when the user wants to benchmark on DLiPath, or asks about evaluating this task. Reports AUC, Accuracy, Precision, Recall, F1-Score.
Evaluates the accuracy of a composable benchmark generation framework in estimating deep learning model inference latency on CPUs, and measures the computational speedup achieved by benchmarking unique layer sequences instead of running full end-to-end models. Use when the user wants to benchmark on 50 DL Models, or asks about evaluating this task. Reports Normalized Latency.
This protocol evaluates information retrieval systems by measuring their ranking effectiveness on passage retrieval tasks using both original seed queries and LLM-generated query variants aligned with specific demographic or textual profiles. It probes whether retrieval systems perform consistently across diverse user personas and query transformations, revealing potential disparities in system behavior and ranking stability. Use when the user wants to benchmark on DL21 & DL22 (TREC Deep Lear...
Evaluates the predictive accuracy and computational efficiency of deep learning models for urban traffic forecasting. It benchmarks grid-based, graph-based, and multivariate time-series architectures on standard traffic datasets to compare their ability to capture spatiotemporal dependencies. Use when the user wants to benchmark on BikeNYC-I, TaxiNYC, TaxiBJ, METR-LA, PeMS-BAY, PEMSD7M, or asks about evaluating this task. Reports MAE.
Evaluates the robustness of deep learning models trained on different frameworks (TensorFlow, Theano, Torch) against adversarial attacks. It measures how well models maintain correct predictions when subjected to white-box, black-box, and decision-based perturbations. Use when the user wants to benchmark on MNIST, CIFAR-10, or asks about evaluating this task. Reports robustness indicator R(m_i).
Evaluates the execution speed and hardware resource utilization of three open-source deep learning frameworks (TensorFlow, Theano, CNTK) across standard computer vision, NLP, and custom datasets. Use when the user wants to benchmark on MNIST, CIFAR-10, IMDB, Self-Driving Car, Penn TreeBank, or asks about evaluating this task. Reports processing_time.