Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,914
skills in category
997
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,001–6,024 of 23,914 skills

Oxford Spires EvalA

Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.

researchpython
0
3
Oversmoothing Rate EvalA

This evaluation probes the tendency of autoregressive neural machine translation models to prematurely terminate sequences by assigning high probability to short prefixes. It measures how well a model balances sequence length distribution and translation quality under beam search decoding. Use when the user wants to benchmark on IWSLT'17, WMT'16 En->De, WMT'19, or asks about evaluating this task. Reports oversmoothing_rate.

researchpython
0
3
Overcooked Ai Adaptation EvalA

This benchmark evaluates the real-time adaptability and communication capabilities of LLM-powered embodied agents in human-robot collaboration. It probes how well agents adjust their high-level subtask planning and low-level movement paths when faced with dynamic, constrained environments and non-adaptive human partners. Use when the user wants to benchmark on Enhanced Overcooked-AI, or asks about evaluating this task. Reports overall score.

researchpythongo
0
3
Oven EvalA

This benchmark probes a model's ability to recognize visual entities in an open-domain setting, specifically testing zero-shot generalization to entities not seen during training. It evaluates how well a model can align visual inputs with structured knowledge graph descriptions to perform entity retrieval. Use when the user wants to benchmark on OVEN, or asks about evaluating this task. Reports Harmonic Mean (HM) of top-1 accuracy.

researchpythontesting
0
3
Ovarian Cancer Subtype EvalA

Evaluates histopathology foundation models and ImageNet-pretrained encoders on classifying ovarian cancer subtypes from whole slide images. It probes the ability of vision models to extract diagnostically relevant features from medical histology slides for multi-class classification. Use when the user wants to benchmark on Ovarian Cancer WSI Dataset, or asks about evaluating this task. Reports balanced accuracy.

researchpythongit
0
3
Ov Vsggen EvalA

Probes a model's ability to generate open-vocabulary video scene graphs by predicting objects, attributes, relations, and triplets from ground-truth trajectories. It evaluates semantic matching accuracy beyond exact label overlap and tests temporal grounding for dynamic relations. Use when the user wants to benchmark on PVSG, VidOR, VIPSeg, SVG2 test set, or asks about evaluating this task. Reports object/attribute/relation/triplet prediction accuracy (LLM-judged).

researchpythongo
0
3
Ov Vg EvalA

Evaluates a model's ability to localize objects in images based on natural language descriptions without prior exposure to those specific categories. It probes visual-linguistic alignment, handling of novel vocabulary, and robustness to varying object scales and complex scenes. Use when the user wants to benchmark on OV-VG, or asks about evaluating this task. Reports Acc50.

researchpythongo
0
3
Outlier Detection EvalA

Evaluates the ability of an outlier detection algorithm to identify anomalous data points in highly imbalanced datasets without prior knowledge of fraud patterns. It probes consistency estimation and ensemble clustering robustness across varying feature spaces and class distributions. Use when the user wants to benchmark on Satimage-2, Thyroid, Credit Card Fraud Detection, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Ott Qa EvalA

This benchmark evaluates a model's ability to perform open-domain question answering by retrieving and fusing evidence from both tabular and textual sources. It specifically probes multi-hop reasoning capabilities where answers require bridging information across separate table segments and text passages. Use when the user wants to benchmark on OTT-QA, or asks about evaluating this task. Reports EM.

researchpythongo
0
3
Otb99 EvalA

Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.

researchpythontesting
0
3
Ota Firmware Update EvalA

Evaluates the energy efficiency and update latency of Over-The-Air (OTA) firmware update strategies on flash-based, batteryless IoT devices under simulated energy-harvesting conditions. Use when the user wants to benchmark on OTA Firmware Update Benchmarks, or asks about evaluating this task. Reports Total Update Energy Consumption.

researchpython
0
3
Ot Detection EvalA

Evaluates a machine learning model's ability to detect overshooting tops (OTs) in satellite imagery at a 2 km pixel resolution. It measures how well the model predicts convection/OT presence using physics-informed features derived from visible and infrared channels. Use when the user wants to benchmark on GOES-16 ABI + MRMS Convection Labels, or asks about evaluating this task. Reports hit, correct rejection, false alarm, miss counts.

researchpythongo
0
3
Osworld Verified EvalA

Evaluates an agent's ability to autonomously plan and execute multi-step GUI automation tasks across various desktop applications. It measures the success rate on both in-distribution tasks from the OSWorld-Verified benchmark and out-of-distribution tasks across six distinct Linux applications. Use when the user wants to benchmark on OSWorld-Verified, OOD GUI Benchmark, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Osworld EvalA

Evaluates multimodal agents' ability to perform open-ended, real-world desktop computer tasks across multiple operating systems. It probes GUI grounding, multi-app workflow navigation, and executable action prediction in dynamic, interactive environments. Use when the user wants to benchmark on OSWorld, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Osteosarcoma Histopathology EvalA

Evaluates deep learning models on classifying osteosarcoma histopathology images into non-tumor, non-viable tumor, viable tumor, and non-viable ratio categories without prior segmentation. Probes the model's ability to capture local texture and global spatial patterns for medical image classification. Use when the user wants to benchmark on TCIA Osteosarcoma, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ost Bench EvalA

Evaluates multimodal large language models' ability to perform online spatio-temporal scene understanding and dynamic, agent-centric reasoning. It tests how well models update spatial and temporal knowledge as they incrementally explore environments, retrieve long-term memory, and infer object relationships across sequential observations. Use when the user wants to benchmark on OST-Bench, or asks about evaluating this task. Reports Overall average score.

researchpythongo
0
3
Osmabench EvalA

Evaluates open semantic mapping models' robustness to dynamic indoor lighting and motion conditions. It measures semantic segmentation accuracy and frequency-weighted IoU, alongside LLM-generated visual question answering accuracy on scene graphs. Use when the user wants to benchmark on ReplicaCAD, HM3D, or asks about evaluating this task. Reports mAcc, f-mIoU.

researchpythongo
0
3
Oscbench EvalA

This benchmark probes a text-to-video model's ability to accurately render and maintain object state transformations (e.g., peeling, slicing) over time. It additionally measures semantic adherence to prompts, scene consistency, and overall perceptual quality to diagnose temporal coherence and physical realism in generated videos. Use when the user wants to benchmark on OSCBench, or asks about evaluating this task. Reports state-change accuracy.

researchpython
0
3
Osbench EvalA

Evaluates subject-driven image generation and manipulation capabilities, specifically testing identity consistency, prompt adherence, and background preservation across single- and multi-subject scenarios. Use when the user wants to benchmark on OSBench, or asks about evaluating this task. Reports Overall (Generation).

researchpythontesting
0
3
Orsi Sod EvalA

This benchmark evaluates optical remote sensing salient object detection models by measuring their ability to accurately segment prominent objects from complex, cluttered backgrounds. It probes structural consistency, boundary precision, and error magnitude across varying object scales and scene complexities. Use when the user wants to benchmark on ORSSD, EORSSD, ORSI-4199, or asks about evaluating this task. Reports maximum F-measure ($F_{\beta}^{max}$).

researchpythonperformance
0
3
Orionbench EvalA

Evaluates unsupervised time series anomaly detection pipelines across diverse real-world and synthetic datasets. It measures detection accuracy for both point and segment anomalies while tracking computational efficiency and model stability over continuous benchmarking cycles. Use when the user wants to benchmark on OrionBench (NASA, NAB, Yahoo S5, UCR), or asks about evaluating this task. Reports F1 score.

researchpythonazure
0
3
Orchid EvalA

Evaluates models on target-independent stance detection (3-way classification) and argumentative dialogue summarization (overall and stance-specific). It probes the ability to classify conflicting viewpoints in Chinese debates and generate concise, faithful summaries aligned with specific stances. Use when the user wants to benchmark on OrChiD, or asks about evaluating this task. Reports Accuracy, ROUGE-1 F1.

researchpythongo
0
3
Orca EvalA

Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.

researchpythongo
0
3
Orbit EvalA

Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.

researchpythonperformance
0
3