Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,574
skills in category
983
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,137–5,160 of 23,574 skills

Rexsenovqa EvalA

This benchmark evaluates video-language models on procedure-centric ultrasound understanding, specifically probing dynamic procedural reasoning, causal troubleshooting, and temporal action understanding. It measures how well models interpret visual evidence and reason through medical imaging procedures without relying on audio or static priors. Use when the user wants to benchmark on ReXSonoVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Rexrank EvalA

Evaluates AI models' ability to generate accurate and clinically relevant radiology reports from chest X-ray images. It assesses both linguistic quality and clinical entity extraction/alignment across diverse clinical datasets. Use when the user wants to benchmark on ReXGradient, MIMIC-CXR, IU X-ray, CheXpert Plus, or asks about evaluating this task. Reports 1/RadCliQ-v1.

researchpythongo
0
3
Rexbench EvalA

This benchmark evaluates the ability of LLM-based coding agents to autonomously implement research extensions by modifying existing AI/ML codebases based on domain-expert instructions. It probes complex, multi-step software engineering capabilities, including codebase navigation, hypothesis-driven implementation, and producing executable patches. Use when the user wants to benchmark on REXBench, or asks about evaluating this task. Reports final success rate.

researchpythongit
0
3
Rewardmap EvalA

Evaluates fine-grained visual reasoning and spatial understanding on structured domains like transit maps. It probes the model's ability to follow complex routes, count stops, and verify true/false statements about visual layouts, while also measuring generalization to broader spatial and chart reasoning benchmarks. Use when the user wants to benchmark on ReasonMap, ReasonMap-Plus, SEED-Bench-2-Plus, SpatialEval, V*Bench, HRBench, ChartQA, MMStar, or asks about evaluating this task. Reports W...

researchpythongo
0
3
Rewardbench2 EvalA

RewardBench2 evaluates reward models across six distinct capabilities: focus, math, safety, factuality, precise instruction following, and ties. It probes whether models can correctly rank a single high-quality response against three inferior completions, and specifically tests robustness in domains with multiple equally valid answers. The benchmark uses unseen human prompts to ensure evaluation independence from downstream post-training tests. Use when the user wants to benchmark on RewardBe...

researchpythongo
0
3
Reward Model Benchmarking EvalA

Evaluates reward models on their ability to correctly rank or score LLM-generated responses across diverse domains like safety, mathematics, coding, and instruction following. It probes pairwise preference accuracy, robustness to superficial biases, cross-sample score calibration, and consistency in assigning absolute quality scores. Use when the user wants to benchmark on RewardBench, RM-Bench, PPE, JudgeBench, or asks about evaluating this task. Reports binary choice accuracy.

researchpythongo
0
3
Revqa Spatial Reasoning EvalA

Evaluates the spatial reasoning and logical comprehension capabilities of multimodal large language models (MLLMs) on synthetic, spatially precise images. It probes robustness to negations, logical operators (AND/OR), adversarial object substitutions, and complex spatial relationships. Use when the user wants to benchmark on RevQA, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Revise EvalA

Evaluates generalized speech enhancement and audio-visual speech resynthesis across multiple distortion types (denoising, separation, inpainting, video-to-speech). Measures content intelligibility, audio-video synchronization, perceptual quality, and low-level signal reconstruction. Use when the user wants to benchmark on LRS3, EasyCom, or asks about evaluating this task. Reports WER.

researchpython
0
3
Reviewmt EvalA

Evaluates LLMs on simulating the academic peer review process across multi-turn dialogues. It probes the model's ability to generate relevant paper summaries, write comprehensive reviews, and make accurate acceptance or rejection decisions based on long-context interactions. Use when the user wants to benchmark on ReviewMT, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Reviewer Too EvalA

Evaluates the ability of LLM-based reviewer agents to predict conference acceptance decisions and generate high-quality peer reviews. It probes classification accuracy against human decisions and assesses review quality through LLM-judged pairwise comparisons across multiple dimensions. Use when the user wants to benchmark on ICLR-2k dataset, or asks about evaluating this task. Reports macro-F1 (5-way).

researchpythongo
0
3
Reviewbench EvalA

Evaluates the quality and accuracy of automated peer reviews generated by LLMs. It probes both the semantic alignment of review text against paper-specific rubrics and the precision of predicted numerical ratings and acceptance decisions. Use when the user wants to benchmark on ReviewBench, or asks about evaluating this task. Reports Rubric Overall Score.

researchpythongo
0
3
Review Quality MetricsA

Evaluates the quality of academic peer reviews by measuring how well criticisms are supported by evidence (substantiation), how factually accurate the review's claims are (correctness), and how thoroughly the review covers the paper's contributions (completeness). Use when the user has predictions and gold and needs to compute correctness.

researchpythongo
0
3
Review Quality EvalA

Evaluates the quality of academic peer review reports across different conferences and years using a multi-dimensional framework. It measures how substantive, actionable, and well-grounded reviews are, and tracks whether these qualities decline over time. Use when the user wants to benchmark on Peer Review Campaigns (ICLR, NeurIPS, ACL), or asks about evaluating this task. Reports Q.

researchpythongit
0
3
Retrotrucks EvalA

Evaluates video anomaly detection models on dashcam footage, specifically testing their ability to detect complex traffic anomalies like collisions and skidding in dynamic, real-world driving scenes. It also benchmarks performance against standard pedestrian anomaly detection datasets to highlight challenges posed by moving cameras and contextual anomalies. Use when the user wants to benchmark on RetroTrucks, UCSD Ped1, UCSD Ped2, ShanghaiTech, or asks about evaluating this task. Reports AUC-...

researchpythongo
0
3
Retrosynthesis Solve Rate EvalA

Evaluates an LLM's ability to generate valid chemical synthesis pathways for target molecules given reference routes and iterative feedback. Use when the user wants to benchmark on Pistachio Hard, or asks about evaluating this task. Reports solve rate.

researchpythongo
0
3
Retrofitting Word Vectors EvalA

Evaluates the semantic quality of pre-trained word vectors by measuring performance improvements after applying a graph-based retrofitting method using semantic lexicons. It probes the model's ability to capture lexical relations (e.g., synonymy, hyponymy) and generalizes across different vector training methods, lexicon types, and languages. Use when the user wants to benchmark on MEN-3k, RG-65, WS-353, TOEFL, SYN-REL, SA, MC-30, or asks about evaluating this task. Reports Spearman's correla...

researchpythongo
0
3
Retrieval Robustness EvalA

This evaluation probes how consistently large language models maintain or improve their answer quality when provided with retrieved context, specifically measuring resilience to variations in retrieval size, document order, and the risk of performance degradation compared to non-retrieval baselines. Use when the user wants to benchmark on Wikipedia QA benchmark, or asks about evaluating this task. Reports No-Degradation Rate (NDR).

researchpythongo
0
3
Retrieval Rerank Rag EvalA

Evaluates the effectiveness of sparse and dense retrieval models, re-rankers, and retrieval-augmented generation (RAG) pipelines on open-domain QA, multi-hop QA, and fact verification tasks. Use when the user wants to benchmark on Natural Questions, TriviaQA, HotpotQA, 2WikiMultiHopQA, ArchivalQA, MSMARCO, WebQuestions, PopQA, ChroniclingAmericaQA, or asks about evaluating this task. Reports Top-k accuracy, Exact match (EM).

researchpythongo
0
3
Retinal Vessel Segmentation EvalA

Evaluates a model's ability to segment retinal blood vessels in fundus images, focusing on preserving fine, elongated vascular structures and handling varying image resolutions and pathologies. The protocol tests robustness under strict hyperparameter consistency and standardized data splits. Use when the user wants to benchmark on DRIVE, STARE, CHASE_DB1, HRF, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Reta Benchmark EvalA

Evaluates the accuracy and structural consistency of retinal vascular tree annotations across pixel, vessel segment, and network levels, ensuring topological correctness and geometrical plausibility. Use when the user wants to benchmark on RETA Benchmark, or asks about evaluating this task. Reports multi_stage_annotation.

researchpythongo
0
3
Resyn EvalA

Evaluates the reasoning capabilities of language models on synthetically generated, code-verifiable tasks. It measures performance on both the custom ReSyn dataset and standard reasoning benchmarks using zero-shot generation with specific sampling parameters. Use when the user wants to benchmark on ReSyn, or asks about evaluating this task. Reports mean@4.

researchpythonperformance
0
3
Restc Sbr EvalA

Evaluates a model's ability to perform session-based next-item recommendation by capturing both spatial graph structures and temporal dynamics. It probes how well the model aggregates collaborative filtering signals and session-specific sequences to predict the subsequent item in a user's browsing session. Use when the user wants to benchmark on Tmall, Diginetica, Gowalla, RetailRocket, Nowplaying, LastFM, or asks about evaluating this task. Reports cross-entropy.

researchpythongo
0
3
Respondeoqa EvalA

This benchmark evaluates large language models on bilingual Latin-English question answering across knowledge-based, skill-based (grammar, scansion, literary devices), multihop reasoning, and translation tasks. It probes models' ability to handle classical language morphology, poetic meter analysis, and cross-lingual generation under constrained and unconstrained settings. Use when the user wants to benchmark on RespondeoQA, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Resource Usage Benchmark EvalA

This evaluation protocol measures the computational efficiency and energy consumption of distributed deep learning training runs. It probes how model architecture, dataset, and hardware constraints (GPU count, power caps, clock speeds) affect training speed and resource utilization. Use when the user wants to benchmark on ImageNet, WikiText-103, QM9, or asks about evaluating this task. Reports training speed.

researchpython
0
3