Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,917
skills in category
997
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,025–6,048 of 23,917 skills

Orchid EvalA

Evaluates models on target-independent stance detection (3-way classification) and argumentative dialogue summarization (overall and stance-specific). It probes the ability to classify conflicting viewpoints in Chinese debates and generate concise, faithful summaries aligned with specific stances. Use when the user wants to benchmark on OrChiD, or asks about evaluating this task. Reports Accuracy, ROUGE-1 F1.

researchpythongo
0
3
Orca EvalA

Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.

researchpythongo
0
3
Orbit EvalA

Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.

researchpythonperformance
0
3
Orb EvalA

Evaluates machine reading comprehension models across diverse linguistic phenomena such as coreference resolution, temporal logic, and causal inference. It tests generalization capabilities by applying synthetic out-of-distribution augmentations to questions, ensuring models rely on reading comprehension rather than external information retrieval. Use when the user wants to benchmark on ORB Benchmark (NewsQA, Quoref, DROP, SQuAD 1.1, SQuAD 2.0, ROPES, DuoRC, NarrativeQA), or asks about evalua...

researchpythongo
0
3
Orangesum EvalA

Evaluates abstractive French summarization quality by measuring lexical overlap, semantic similarity, and human judgments of accuracy, informativeness, and fluency. Use when the user wants to benchmark on OrangeSum, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
Orad 3d EvalA

Evaluates off-road autonomous driving capabilities across perception, planning, and world modeling. It probes 2D free-space detection, 3D semantic occupancy prediction, GPS-guided trajectory planning, VLM-based scene understanding and path planning, and future video generation in unstructured, variable-terrain environments. Use when the user wants to benchmark on ORAD-3D, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
Opus Mt EvalA

Evaluates the translation quality of OPUS-MT models across diverse language pairs using standard automatic metrics and specialized linguistic test suites. It probes general-purpose translation capability, lexical ambiguity disambiguation, and cross-lingual robustness. Use when the user wants to benchmark on Flores, Tatoeba, MuCoW, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Opus 100 Nmt EvalA

Evaluates massively multilingual neural machine translation models on translation quality and language accuracy across 100 languages, including zero-shot translation between unseen language pairs. Use when the user wants to benchmark on OPUS-100, or asks about evaluating this task. Reports BLEU_94.

researchpythongit
0
3
Optimam Mammography EvalA

Binary classification of high-resolution mammography images to detect malignant breast tissue. It probes the model's ability to distinguish between malignant and non-malignant cases using both localized patches and full-resolution inputs. Use when the user wants to benchmark on OPTIMAM, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Optimam Mammography Cad EvalA

Evaluates a deep learning object detection and classification system for identifying malignant lesions in mammographic images. It probes the model's ability to handle high-resolution medical imaging data and diverse lesion morphologies (e.g., microcalcifications, masses) using deformable convolutions. Use when the user wants to benchmark on Optimam, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Optiloop 5g Energy EvalA

Evaluates the energy efficiency and operational performance of a 5G network orchestration framework under dynamic traffic conditions. It probes the system's ability to jointly optimize virtual network function (VNF) placement, traffic routing, and network element activation to minimize power consumption while maintaining connectivity and processing capacity. Use when the user wants to benchmark on Real-world mobile operator traffic snapshot, or asks about evaluating this task. Reports energy ...

researchpythonnode
0
3
Optical Flow Kitti Sintel EvalA

This evaluation protocol measures the accuracy of predicted optical flow fields against ground truth displacement vectors across varying motion magnitudes and occlusion conditions. It probes a model's ability to handle both small, fine-grained movements and large, robust displacements using standard video sequence benchmarks. Use when the user wants to benchmark on KITTI2012, KITTI2015, MPI-Sintel, or asks about evaluating this task. Reports Out-Noc, EPE.

researchpythonperformance
0
3
Optical Flow EvalA

Evaluates dense optical flow estimation by predicting pixel-wise displacement vectors between consecutive frames. It probes robustness to large motions, occlusions, blur, and atmospheric effects across synthetic and real-world driving scenes. Use when the user wants to benchmark on MPI Sintel, KITTI, Middlebury, or asks about evaluating this task. Reports AEE (Average Endpoint Error).

researchpython
0
3
Optical Flow Estimation EvalA

Evaluates the accuracy of predicted optical flow fields against ground truth motion vectors between consecutive image frames. It probes a model's ability to estimate dense pixel-wise displacement in both synthetic cinematic scenes and real-world driving environments. Use when the user wants to benchmark on FlyingChairs, Sintel, KITTI12, KITTI15, Middlebury, or asks about evaluating this task. Reports AEE.

researchpython
0
3
Optical Flow Epe EvalA

Evaluates the accuracy of predicted optical flow fields against ground truth displacements between consecutive video frames. It probes a model's ability to handle occlusions, non-rigid motion, and large displacements in both synthetic and real-world driving scenarios. Use when the user wants to benchmark on Sintel, KITTI 2012, or asks about evaluating this task. Reports end point error.

researchpython
0
3
Optbench EvalA

Evaluates formal theorem proving capabilities specifically within the undergraduate optimization domain. It probes a model's ability to generate syntactically correct and semantically progressive Lean 4 proof steps or full scripts under strict verifier constraints, while measuring robustness against catastrophic forgetting on general math benchmarks. Use when the user wants to benchmark on OptBench, MiniF2F-test, ProofNet-test, or asks about evaluating this task. Reports Pass@32.

researchpythongo
0
3
Opt Iml Bench EvalA

Evaluates instruction-tuned language models on generalization across 1,991 NLP tasks spanning 100+ categories. It probes zero-shot and few-shot (5-shot) performance on held-out categories, unseen tasks within seen categories, and fully supervised tasks, measuring both generation quality and classification accuracy. Use when the user wants to benchmark on OPT-IML Bench, or asks about evaluating this task. Reports Rouge-L.

researchpythongo
0
3
Opinion Summarization EvalA

Evaluates a model's ability to synthesize multiple product reviews into a single, coherent opinion summary. It probes the model's capacity to capture diverse aspects, maintain factual accuracy, and adhere to domain-specific stylistic constraints without relying on surface-level lexical overlap. Use when the user wants to benchmark on Amazon, Oposum+, Flipkart, or asks about evaluating this task. Reports human evaluation.

researchpythongo
0
3
Ophthalmic Multimodal EvalA

Evaluates multimodal large language models on clinical ophthalmic image interpretation, specifically diagnosing retinal and macular diseases from fundus photographs and optical coherence tomography (OCT) scans. It probes the models' ability to recognize normal conditions and identify specific pathological states across diverse disease categories. Use when the user wants to benchmark on Ophthalmic Multimodal Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ophnet EvalA

Evaluates video understanding models on ophthalmic surgical workflows, including recognizing primary surgery types, classifying hierarchical surgical phases and operations, localizing phase boundaries in time, and anticipating upcoming surgical phases. Use when the user wants to benchmark on OphNet, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Openxai EvalA

Evaluates the faithfulness, stability, and fairness of post-hoc feature attribution explanation methods (e.g., LIME, SHAP, gradient-based) on tabular datasets to enable reproducible and transparent comparisons. Use when the user wants to benchmark on Popular tabular datasets for XAI and fairness research, or asks about evaluating this task. Reports faithfulness.

researchpythongo
0
3
Openvlthinkerv2 EvalA

Evaluates a multimodal reasoning model's capability across diverse visual tasks, including general and mathematical VQA, document understanding, spatial reasoning, and visual grounding. The protocol tests the model's ability to balance fine-grained perception with multi-step reasoning under a unified RL training framework. Use when the user wants to benchmark on MMMU, MMBench, MMStar, ChartQA, DocVQA, OCRBench, InfoVQA, EmbSpatial, RefSpatial, RoboSpatial, RefCOCO, RefCOCO+, RefCOCOg, or asks...

researchpythongo
0
3
Openvid 1m EvalA

Evaluates text-to-video generation models on visual aesthetics, technical quality, text-video alignment, and temporal consistency using a standardized set of 700 prompts. Use when the user wants to benchmark on Liu et al. (2023b) Benchmark, or asks about evaluating this task. Reports VQAA.

researchpythongo
0
3
Openve Bench EvalA

Evaluates instruction-guided video editing models on spatially-aligned and non-spatially-aligned editing tasks, probing temporal consistency, spatial fidelity, and instruction following. Use when the user wants to benchmark on OpenVE-Bench, or asks about evaluating this task. Reports overall score.

researchpythongo
0
3