Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,574
skills in category
983
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 5,065–5,088 of 23,574 skills

Rover EvalA

Evaluates the accuracy and robustness of visual-inertial SLAM systems across diverse outdoor environments, seasons, and lighting conditions. It probes long-term trajectory consistency, scale estimation, and environmental adaptability under challenging visual degradation. Use when the user wants to benchmark on ROVER, or asks about evaluating this task. Reports mATE.

researchpythonperformance
0
3
Routing Algorithm Benchmarks EvalA

Evaluates the routing algorithm's efficiency, scalability, and classification accuracy on standard NLP and vision benchmarks when used as a classification head over frozen pretrained Transformers. Use when the user wants to benchmark on IMDB, SST-5, SST-2, ImageNet-1K, CIFAR-100, CIFAR-10, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Routenlp EvalA

Evaluates a closed-loop LLM routing system's ability to dynamically select between a four-tier model portfolio based on task difficulty, balancing inference cost, response quality, and latency. The benchmark probes how well a router can escalate queries to more capable models only when necessary, while using distillation and conformal cascading to maintain performance at lower cost tiers. Use when the user wants to benchmark on EDGAR (NER), EDGAR (Summarization), BANKING77* (Intent Classifica...

researchpythongo
0
3
Routefinder Vrp EvalA

Evaluates a foundation model's ability to solve and generalize across 48 distinct Vehicle Routing Problem (VRP) variants. It probes constraint satisfaction, route optimization, and zero-shot adaptation to unseen attribute combinations like multi-depots and mixed backhauls. Use when the user wants to benchmark on VRP variants (unified generation), or asks about evaluating this task. Reports optimality gap.

researchpythongit
0
3
Roundtable Policy EvalA

Evaluates LLMs' scientific reasoning and proposal writing capabilities through a multi-task accuracy benchmark and a rubric-based narrative generation task. It also assesses the stability of consensus methods and the consistency of AI graders across structured and open-ended scientific domains. Use when the user wants to benchmark on MultiTask scientific tasks, SingleTask scientific proposal writing, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Roundabout Tau EvalA

This benchmark evaluates a model's ability to detect traffic anomalies in roadside surveillance videos and generate detailed, reasoning-grounded textual summaries of those anomalies. It probes both binary/fine-grained classification accuracy and semantic alignment of generated descriptions with ground-truth event narratives. Use when the user wants to benchmark on Roundabout-TAU, or asks about evaluating this task. Reports 4-cls AP.

researchpythongo
0
3
Roughness Index DistanceA

Evaluates surface irregularities and topological consistency between predicted and ground-truth 3D medical segmentation masks. It quantifies local surface roughness, relative roughness differences, and average surface distance to detect spikes, holes, and smoothing artifacts. Use when the user has predictions and gold and needs to compute Roughness Index (RI).

researchpythongo
0
3
Rood Mri EvalA

Evaluates the robustness of deep learning segmentation models to out-of-distribution MRI data and synthetic corruptions (noise, contrast, resolution, spatial shifts, motion artifacts) across multiple severity levels. It measures performance degradation on anatomical and lesion segmentation tasks compared to clean data. Use when the user wants to benchmark on ROOD-MRI Benchmark (Hippocampus, Ventricle, WMH), or asks about evaluating this task. Reports DSC.

researchpythonexpress
0
3
Romeo Vuln Detection EvalA

This benchmark evaluates binary vulnerability detection capabilities on assembly language representations of C/C++ functions. It probes whether models can identify security flaws (e.g., buffer overflows, integer overflows) by analyzing machine code semantics and call graph context. Use when the user wants to benchmark on ROMEO, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Romath EvalA

This benchmark evaluates large language models' mathematical reasoning capabilities specifically in Romanian. It probes the ability to solve single-step and multi-step problems, handle verifiable numerical answers, and construct or verify mathematical proofs without relying on direct English translations. Use when the user wants to benchmark on RoMath, or asks about evaluating this task. Reports correctness.

researchpythongo
0
3
Rodeo Reconstruction EvalA

Evaluates a deep learning autoencoder's ability to reconstruct high-quality images from undersampled or noisy data across synthetic, MRI, and CT domains. It probes robustness to impulse noise, Fourier undersampling, and sparse tomographic projections compared to compressed sensing and standard autoencoders. Use when the user wants to benchmark on CIFAR-10, Cardiac Perfusion MRI, Larynx & Cardiac MRI, Speech MRI, ULB CT Dataset, or asks about evaluating this task. Reports NMSE.

researchpythongo
0
3
Rodent Bench EvalA

Evaluates multimodal large language models on temporal segmentation and fine-grained behavioral annotation of rodent videos. It probes capabilities in long-video processing, distinguishing subtle or rare behaviors, and handling diverse experimental paradigms and camera angles. Use when the user wants to benchmark on Rodent-Bench-Long, Rodent-Bench-Short, or asks about evaluating this task. Reports Weighted Matthew’s Correlation Coefficient (MCC).

researchpythonperformance
0
3
Rococo EvalA

Evaluates the robustness of image-text matching models against adversarial perturbations injected into the retrieval gallery. It probes whether models rely on holistic semantic alignment or are easily misled by locally similar but semantically altered images and captions. Use when the user wants to benchmark on MS-COCO (RoCOCO variant), or asks about evaluating this task. Reports Recall@1.

researchpythongo
0
3
Robustspring EvalA

Evaluates the robustness of dense correspondence models (optical flow, scene flow, stereo) to 20 types of image corruptions by measuring the divergence between predictions on clean images and predictions on corrupted images. Use when the user wants to benchmark on Spring, or asks about evaluating this task. Reports R^c_EPE.

researchpythonspring
0
3
Robusto 1 EvalA

Evaluates the cognitive alignment and visuocognitive reasoning of Vision-Language Models (VLMs) compared to humans on real-world, out-of-distribution autonomous driving scenarios from Peru. It probes how models and humans interpret complex, rare driving situations through open-ended, multiple-choice, and counterfactual/hypothetical visual question answering. Use when the user wants to benchmark on Robusto-1, or asks about evaluating this task. Reports Representational Similarity Analysis (RSA).

researchpython
0
3
Robustness Perturbation EvalA

Evaluates the robustness of neural language models to non-adversarial character- and word-level input perturbations (e.g., typos, deletions, synonyms) while preserving semantic meaning. Use when the user wants to benchmark on TC, SA, NER, SS, QA (unspecified downstream datasets), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Robustbench Cifar10 EvalA

Evaluates adversarial robustness of image classifiers under multi-norm threat models, testing whether pre-screening diagnostics (FOSC, RDI) reliably predict full attack performance and expose worst-case vulnerabilities masked by single-norm evaluations. Use when the user wants to benchmark on RobustBench CIFAR-10, or asks about evaluating this task. Reports robust_accuracy.

researchpythongo
0
3
Robust Mvd EvalA

Evaluates multi-view depth estimation models on their ability to generalize across diverse domains, scales, and camera configurations. It specifically probes robustness to out-of-distribution cost volume statistics and tests absolute scale estimation without requiring scale alignment or depth range assumptions. Use when the user wants to benchmark on StaticThings3D, BlendedMVS, or asks about evaluating this task. Reports depth error metrics (e.g., RMSE, AbsRel, δ1).

researchpython
0
3
Robust Mean Relative ErrorA

Evaluates a model's ability to forecast future values in a financial time-series dataset given historical observations. It measures prediction accuracy using a capped relative error to prevent extreme outliers from dominating the score. Use when the user has predictions and gold and needs to compute robust mean relative error.

researchpythongo
0
3
Robust Asr Wer EvalA

Evaluates the robustness of end-to-end automatic speech recognition models against various stationary and non-stationary noise types at different signal-to-noise ratios (SNR). It also measures the degradation of recognition accuracy on clean speech when noise-adaptation techniques are applied. Use when the user wants to benchmark on Custom noisy speech dataset (7 noise types), or asks about evaluating this task. Reports WER.

researchpythongit
0
3
Robust Asr EvalA

Evaluates the robustness of monaural automatic speech recognition systems under noisy and reverberant conditions. It measures how well a decoupled frontend speech enhancement module improves the word error rate of a backend ASR model trained exclusively on clean speech. Use when the user wants to benchmark on WSJ0 SI-84, CHiME-2, LibriSpeech, or asks about evaluating this task. Reports WER.

researchpythonfrontend
0
3
Roboverse Imitation EvalA

Evaluates robot manipulation policies on a unified set of contact-rich pick-and-place and articulation tasks across multiple simulators. It probes both specialist and generalist vision-language-action models on success rates under standard and progressively challenging generalization levels. Use when the user wants to benchmark on ROBOVERSE Imitation Learning Benchmark, or asks about evaluating this task. Reports success rate.

researchpythontesting
0
3
Robotracer Spatial EvalA

Evaluates vision-language models' ability to perform spatial understanding, metric measuring, 2D/3D referring, and multi-step visual tracing in cluttered environments. It probes geometric reasoning, depth estimation, and collision-free path planning for robotic manipulation. Use when the user wants to benchmark on CV-Bench, BLINK_val, RoboSpatial, Embspacial, Q-spatial, MSMU, Where2Place, RefSpatial-Bench, ShareRobot-Bench, VABench-V, TraceSpatial-Bench, RoboTwin, MMEtest, MMBenchdev, OK-VQA,...

researchpythongo
0
3
Robotic Manipulation EvalA

Evaluates the ability of diffusion transformer policies to perform long-horizon robotic manipulation tasks across bi-manual, single-arm, and simulated environments. It probes stable training, observation tokenization, and generalization across different robot morphologies and action spaces. Use when the user wants to benchmark on Robotic Manipulation Task Suite, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3