Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,137–5,160 of 23,574 skills
This benchmark evaluates video-language models on procedure-centric ultrasound understanding, specifically probing dynamic procedural reasoning, causal troubleshooting, and temporal action understanding. It measures how well models interpret visual evidence and reason through medical imaging procedures without relying on audio or static priors. Use when the user wants to benchmark on ReXSonoVQA, or asks about evaluating this task. Reports accuracy.
Evaluates AI models' ability to generate accurate and clinically relevant radiology reports from chest X-ray images. It assesses both linguistic quality and clinical entity extraction/alignment across diverse clinical datasets. Use when the user wants to benchmark on ReXGradient, MIMIC-CXR, IU X-ray, CheXpert Plus, or asks about evaluating this task. Reports 1/RadCliQ-v1.
This benchmark evaluates the ability of LLM-based coding agents to autonomously implement research extensions by modifying existing AI/ML codebases based on domain-expert instructions. It probes complex, multi-step software engineering capabilities, including codebase navigation, hypothesis-driven implementation, and producing executable patches. Use when the user wants to benchmark on REXBench, or asks about evaluating this task. Reports final success rate.
Evaluates fine-grained visual reasoning and spatial understanding on structured domains like transit maps. It probes the model's ability to follow complex routes, count stops, and verify true/false statements about visual layouts, while also measuring generalization to broader spatial and chart reasoning benchmarks. Use when the user wants to benchmark on ReasonMap, ReasonMap-Plus, SEED-Bench-2-Plus, SpatialEval, V*Bench, HRBench, ChartQA, MMStar, or asks about evaluating this task. Reports W...
RewardBench2 evaluates reward models across six distinct capabilities: focus, math, safety, factuality, precise instruction following, and ties. It probes whether models can correctly rank a single high-quality response against three inferior completions, and specifically tests robustness in domains with multiple equally valid answers. The benchmark uses unseen human prompts to ensure evaluation independence from downstream post-training tests. Use when the user wants to benchmark on RewardBe...
Evaluates reward models on their ability to correctly rank or score LLM-generated responses across diverse domains like safety, mathematics, coding, and instruction following. It probes pairwise preference accuracy, robustness to superficial biases, cross-sample score calibration, and consistency in assigning absolute quality scores. Use when the user wants to benchmark on RewardBench, RM-Bench, PPE, JudgeBench, or asks about evaluating this task. Reports binary choice accuracy.
Evaluates the spatial reasoning and logical comprehension capabilities of multimodal large language models (MLLMs) on synthetic, spatially precise images. It probes robustness to negations, logical operators (AND/OR), adversarial object substitutions, and complex spatial relationships. Use when the user wants to benchmark on RevQA, or asks about evaluating this task. Reports performance.
Evaluates generalized speech enhancement and audio-visual speech resynthesis across multiple distortion types (denoising, separation, inpainting, video-to-speech). Measures content intelligibility, audio-video synchronization, perceptual quality, and low-level signal reconstruction. Use when the user wants to benchmark on LRS3, EasyCom, or asks about evaluating this task. Reports WER.
Evaluates LLMs on simulating the academic peer review process across multi-turn dialogues. It probes the model's ability to generate relevant paper summaries, write comprehensive reviews, and make accurate acceptance or rejection decisions based on long-context interactions. Use when the user wants to benchmark on ReviewMT, or asks about evaluating this task. Reports F1-score.
Evaluates the ability of LLM-based reviewer agents to predict conference acceptance decisions and generate high-quality peer reviews. It probes classification accuracy against human decisions and assesses review quality through LLM-judged pairwise comparisons across multiple dimensions. Use when the user wants to benchmark on ICLR-2k dataset, or asks about evaluating this task. Reports macro-F1 (5-way).
Evaluates the quality and accuracy of automated peer reviews generated by LLMs. It probes both the semantic alignment of review text against paper-specific rubrics and the precision of predicted numerical ratings and acceptance decisions. Use when the user wants to benchmark on ReviewBench, or asks about evaluating this task. Reports Rubric Overall Score.
Evaluates the quality of academic peer reviews by measuring how well criticisms are supported by evidence (substantiation), how factually accurate the review's claims are (correctness), and how thoroughly the review covers the paper's contributions (completeness). Use when the user has predictions and gold and needs to compute correctness.
Evaluates the quality of academic peer review reports across different conferences and years using a multi-dimensional framework. It measures how substantive, actionable, and well-grounded reviews are, and tracks whether these qualities decline over time. Use when the user wants to benchmark on Peer Review Campaigns (ICLR, NeurIPS, ACL), or asks about evaluating this task. Reports Q.
Evaluates video anomaly detection models on dashcam footage, specifically testing their ability to detect complex traffic anomalies like collisions and skidding in dynamic, real-world driving scenes. It also benchmarks performance against standard pedestrian anomaly detection datasets to highlight challenges posed by moving cameras and contextual anomalies. Use when the user wants to benchmark on RetroTrucks, UCSD Ped1, UCSD Ped2, ShanghaiTech, or asks about evaluating this task. Reports AUC-...
Evaluates an LLM's ability to generate valid chemical synthesis pathways for target molecules given reference routes and iterative feedback. Use when the user wants to benchmark on Pistachio Hard, or asks about evaluating this task. Reports solve rate.
Evaluates the semantic quality of pre-trained word vectors by measuring performance improvements after applying a graph-based retrofitting method using semantic lexicons. It probes the model's ability to capture lexical relations (e.g., synonymy, hyponymy) and generalizes across different vector training methods, lexicon types, and languages. Use when the user wants to benchmark on MEN-3k, RG-65, WS-353, TOEFL, SYN-REL, SA, MC-30, or asks about evaluating this task. Reports Spearman's correla...
This evaluation probes how consistently large language models maintain or improve their answer quality when provided with retrieved context, specifically measuring resilience to variations in retrieval size, document order, and the risk of performance degradation compared to non-retrieval baselines. Use when the user wants to benchmark on Wikipedia QA benchmark, or asks about evaluating this task. Reports No-Degradation Rate (NDR).
Evaluates the effectiveness of sparse and dense retrieval models, re-rankers, and retrieval-augmented generation (RAG) pipelines on open-domain QA, multi-hop QA, and fact verification tasks. Use when the user wants to benchmark on Natural Questions, TriviaQA, HotpotQA, 2WikiMultiHopQA, ArchivalQA, MSMARCO, WebQuestions, PopQA, ChroniclingAmericaQA, or asks about evaluating this task. Reports Top-k accuracy, Exact match (EM).
Evaluates a model's ability to segment retinal blood vessels in fundus images, focusing on preserving fine, elongated vascular structures and handling varying image resolutions and pathologies. The protocol tests robustness under strict hyperparameter consistency and standardized data splits. Use when the user wants to benchmark on DRIVE, STARE, CHASE_DB1, HRF, or asks about evaluating this task. Reports F1 score.
Evaluates the accuracy and structural consistency of retinal vascular tree annotations across pixel, vessel segment, and network levels, ensuring topological correctness and geometrical plausibility. Use when the user wants to benchmark on RETA Benchmark, or asks about evaluating this task. Reports multi_stage_annotation.
Evaluates the reasoning capabilities of language models on synthetically generated, code-verifiable tasks. It measures performance on both the custom ReSyn dataset and standard reasoning benchmarks using zero-shot generation with specific sampling parameters. Use when the user wants to benchmark on ReSyn, or asks about evaluating this task. Reports mean@4.
Evaluates a model's ability to perform session-based next-item recommendation by capturing both spatial graph structures and temporal dynamics. It probes how well the model aggregates collaborative filtering signals and session-specific sequences to predict the subsequent item in a user's browsing session. Use when the user wants to benchmark on Tmall, Diginetica, Gowalla, RetailRocket, Nowplaying, LastFM, or asks about evaluating this task. Reports cross-entropy.
This benchmark evaluates large language models on bilingual Latin-English question answering across knowledge-based, skill-based (grammar, scansion, literary devices), multihop reasoning, and translation tasks. It probes models' ability to handle classical language morphology, poetic meter analysis, and cross-lingual generation under constrained and unconstrained settings. Use when the user wants to benchmark on RespondeoQA, or asks about evaluating this task. Reports exact-match accuracy.
This evaluation protocol measures the computational efficiency and energy consumption of distributed deep learning training runs. It probes how model architecture, dataset, and hardware constraints (GPU count, power caps, clock speeds) affect training speed and resource utilization. Use when the user wants to benchmark on ImageNet, WikiText-103, QM9, or asks about evaluating this task. Reports training speed.