Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,921–4,944 of 23,576 skills
This evaluation protocol measures the impact of score ties on document ranking repeatability across diverse information retrieval collections. It quantifies how non-deterministic tie-breaking during multi-threaded indexing causes variability in standard ranking metrics, even when using identical queries and ranking models. Use when the user wants to benchmark on TREC 2004 Robust Track (Disks 4 & 5), TREC 2005 Robust Track (AQUAINT), TREC 2017 Common Core Track (NYT Annotated Corpus), TREC 201...
Evaluates dense optical flow estimation accuracy and occlusion detection on standard video benchmarks. Probes the model's ability to predict pixel-wise motion vectors and identify occluded regions under varying motion magnitudes and scene complexities. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports End Point Error (EPE).
This evaluation protocol assesses the ability of deep learning models to predict drug-target interactions (DTI) by learning from molecular graphs and protein sequences. It probes the model's capacity to capture cross-domain interaction patterns between small molecules and proteins, particularly in semi-inductive settings where novel compounds are paired with known protein families. Use when the user wants to benchmark on BindingDB, KIBA, Human, SCOPE, or asks about evaluating this task. Repor...
Evaluates computational methods for single-cell multi-omics integration by measuring their ability to preserve biological variation, align different omics layers at the cell and single-cell levels, and enable accurate cell type annotation. Use when the user wants to benchmark on SHARE-seq BMMC, SNARE-seq, 10X Genomics Multiome, Human fetal atlas, CITE-seq BMMC S1, CITE-seq BMMC S4, Human brain multi-omics, Human brain 3k, or asks about evaluating this task. Reports biological variation conser...
Evaluates the fidelity of a particle filter algorithm for generating counterfactual samples from structural causal models by comparing empirical statistics of the generated samples against known ground-truth distributions and correlations. Use when the user wants to benchmark on Synthetic SCM Simulation, or asks about evaluating this task. Reports proportion of unique observations.
Evaluates a model's ability to perform hierarchical scientific summarization by generating three distinct granularity levels (Abstract, Key Contributions, TL;DR) from a single full-text input. It probes multi-granularity text compression and the model's capacity to maintain coherence across varying compression ratios within a single inference pass. Use when the user wants to benchmark on SciZoom, or asks about evaluating this task. Reports unspecified summarization metric.
Evaluates multimodal LLMs on closed-ended visual and non-visual question answering over scientific figures. It probes recognition of visual attributes (color, shape, position) and reasoning capabilities across diverse chart types. Use when the user wants to benchmark on SciVQA, or asks about evaluating this task. Reports ROUGE-1 F1.
Evaluates large language models across four dimensions of trustworthiness in scientific contexts: truthfulness, adversarial robustness, scientific safety, and scientific ethics. It probes models' ability to provide accurate scientific information, resist adversarial perturbations, avoid generating harmful content, and make sound ethical judgments in research scenarios. Use when the user wants to benchmark on SciQ, ARC-C, MMLU, GPQA-Diamond, LogiQA, ReClor, LOGICINFERENCE, WMDP, HarmBench, Sci...
Evaluates long-context language models' ability to perform numerical aggregation, filtering, sorting, and logical operations across extended contexts (up to 1M tokens) using scientific article metadata and full-text articles. Use when the user wants to benchmark on SciTrek, or asks about evaluating this task. Reports exact match.
This benchmark evaluates the ability of models to generate extreme, single-sentence summaries (TLDRs) of scientific papers, capturing key contributions while bypassing background details. It tests both automated overlap metrics and human-judged informativeness and correctness under multi-target and multi-input settings. Use when the user wants to benchmark on SCITLDR, or asks about evaluating this task. Reports Rouge-1.
Evaluates an LLM's ability to comprehend algorithmic descriptions from academic papers and translate them into executable code. It probes the model's capacity for algorithmic reasoning, dependency resolution, and practical implementation within a repository context. Use when the user wants to benchmark on SciReplicate-Bench, or asks about evaluating this task. Reports Execution Accuracy.
Evaluates open-ended, closed-book scientific question answering capabilities. It probes a model's ability to generate comprehensive, accurate, and reasonable answers to research-level science questions without external context or reference papers. Use when the user wants to benchmark on SciQAG-24D, SciQ, or asks about evaluating this task. Reports CAR.
Evaluates a model's ability to perform complex, claim-centric reasoning over full scientific documents containing multimodal elements (charts, tables, figures). It probes the model's capacity to localize evidence and answer questions accurately despite long-context noise and distractors. Use when the user wants to benchmark on SciMDR-Eval, or asks about evaluating this task. Reports accuracy.
Probes large language models' scientific knowledge across five progressive cognitive levels: memory, comprehension, reasoning, ethical discernment, and real-world application. Covers four scientific domains (biology, chemistry, physics, materials science) using diverse question formats including multiple-choice, relation extraction, and open-ended protocol design. Use when the user wants to benchmark on SciKnowEval, or asks about evaluating this task. Reports overall normalized score.
Evaluates multi-modal large language models' ability to interpret scientific graphs and generate accurate, context-aware answers in a multi-turn conversational setting. It probes open-vocabulary visual reasoning and the model's capacity to leverage auxiliary paper metadata for grounded responses. Use when the user wants to benchmark on SciGraphQA, or asks about evaluating this task. Reports CIDEr.
Evaluates the logical correctness, structural fidelity, and information utility of AI-generated scientific images. It probes whether generated visuals accurately encode domain-specific facts and geometric relationships, and whether they are indispensable for solving visually grounded scientific quizzes. Use when the user wants to benchmark on SciGenBench, or asks about evaluating this task. Reports inverse_validation_rate.
This benchmark evaluates large multimodal models' ability to interpret scientific figures by testing their capacity to match figures to captions and vice versa. It probes fine-grained visual-textual reasoning, attention to scientific details, and robustness against adversarially selected distractors. Use when the user wants to benchmark on SciFIBench, or asks about evaluating this task. Reports accuracy.
Evaluates large language models on solving university-level scientific exams in computer science. It probes capabilities in open-ended reasoning, mathematical proof writing, long-form explanations, and multimodal (image-text) understanding across English and German languages. Use when the user wants to benchmark on SciEx, or asks about evaluating this task. Reports Normalized score (0-100%).
Evaluates large language models' scientific intelligence across seven core dimensions, including multimodal perception, understanding, reasoning, knowledge comprehension, code generation, symbolic reasoning, and hypothesis generation. It covers multiple scientific disciplines using both text-only and multimodal inputs to assess real-world scientific workflow capabilities. Use when the user wants to benchmark on SLAKE, MSEarth, SFE, OmniEarth, OmniMedVQA, PhyX, ChemBench, ChemBench4K, LLM4Chem...
Evaluates a model's ability to automatically assess K-12 science instructional materials against pedagogical rubrics. It probes domain-aligned reasoning, long-context evidence grounding, and the capacity to generate rubric-consistent scores and justifications. Use when the user wants to benchmark on SciEval, or asks about evaluating this task. Reports Evidence Match Rate (EMR).
Evaluates the ability of language models to accurately classify scientific abstracts into fine-grained disciplinary or sub-disciplinary categories under few-shot and zero-shot conditions. Use when the user wants to benchmark on SDPRA 2021, arXiv, S2ORC, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness and cross-dataset/domain generalization of relation extraction models on scientific abstracts. It probes how annotation discrepancies and domain shifts affect relation classification performance. Use when the user wants to benchmark on SemEval-2018, SciERC, or asks about evaluating this task. Reports Macro F1-score.
Evaluates the ability of LLMs to generate novel, feasible, and effective scientific research ideas given a research question. Probes open-ended scientific reasoning and ideation quality under compute-matched inference budgets. Use when the user wants to benchmark on ICLR 2024 & NeurIPS 2025, or asks about evaluating this task. Reports Absolute Novelty.
Evaluates a model's ability to perform high-level visual reasoning and domain-specific knowledge grounding on scientific figures within a multiple-choice question answering setting. It specifically probes whether models can resist choice-induced prior bias where text-only answer options incorrectly steer predictions away from visually supported ground truth. Use when the user wants to benchmark on MAC, SciFIBench, MMSci, or asks about evaluating this task. Reports Accuracy.