Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,001–6,024 of 23,914 skills
Evaluates large-scale outdoor localization, 3D reconstruction, and novel-view synthesis using synchronized LiDAR, visual, and IMU data against millimetre-accurate TLS ground truth. Use when the user wants to benchmark on Oxford Spires Dataset, or asks about evaluating this task. Reports metric ground truth.
This evaluation probes the tendency of autoregressive neural machine translation models to prematurely terminate sequences by assigning high probability to short prefixes. It measures how well a model balances sequence length distribution and translation quality under beam search decoding. Use when the user wants to benchmark on IWSLT'17, WMT'16 En->De, WMT'19, or asks about evaluating this task. Reports oversmoothing_rate.
This benchmark evaluates the real-time adaptability and communication capabilities of LLM-powered embodied agents in human-robot collaboration. It probes how well agents adjust their high-level subtask planning and low-level movement paths when faced with dynamic, constrained environments and non-adaptive human partners. Use when the user wants to benchmark on Enhanced Overcooked-AI, or asks about evaluating this task. Reports overall score.
This benchmark probes a model's ability to recognize visual entities in an open-domain setting, specifically testing zero-shot generalization to entities not seen during training. It evaluates how well a model can align visual inputs with structured knowledge graph descriptions to perform entity retrieval. Use when the user wants to benchmark on OVEN, or asks about evaluating this task. Reports Harmonic Mean (HM) of top-1 accuracy.
Evaluates histopathology foundation models and ImageNet-pretrained encoders on classifying ovarian cancer subtypes from whole slide images. It probes the ability of vision models to extract diagnostically relevant features from medical histology slides for multi-class classification. Use when the user wants to benchmark on Ovarian Cancer WSI Dataset, or asks about evaluating this task. Reports balanced accuracy.
Probes a model's ability to generate open-vocabulary video scene graphs by predicting objects, attributes, relations, and triplets from ground-truth trajectories. It evaluates semantic matching accuracy beyond exact label overlap and tests temporal grounding for dynamic relations. Use when the user wants to benchmark on PVSG, VidOR, VIPSeg, SVG2 test set, or asks about evaluating this task. Reports object/attribute/relation/triplet prediction accuracy (LLM-judged).
Evaluates a model's ability to localize objects in images based on natural language descriptions without prior exposure to those specific categories. It probes visual-linguistic alignment, handling of novel vocabulary, and robustness to varying object scales and complex scenes. Use when the user wants to benchmark on OV-VG, or asks about evaluating this task. Reports Acc50.
Evaluates the ability of an outlier detection algorithm to identify anomalous data points in highly imbalanced datasets without prior knowledge of fraud patterns. It probes consistency estimation and ensemble clustering robustness across varying feature spaces and class distributions. Use when the user wants to benchmark on Satimage-2, Thyroid, Credit Card Fraud Detection, or asks about evaluating this task. Reports AUPRC.
This benchmark evaluates a model's ability to perform open-domain question answering by retrieving and fusing evidence from both tabular and textual sources. It specifically probes multi-hop reasoning capabilities where answers require bridging information across separate table segments and text passages. Use when the user wants to benchmark on OTT-QA, or asks about evaluating this task. Reports EM.
Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
Evaluates the energy efficiency and update latency of Over-The-Air (OTA) firmware update strategies on flash-based, batteryless IoT devices under simulated energy-harvesting conditions. Use when the user wants to benchmark on OTA Firmware Update Benchmarks, or asks about evaluating this task. Reports Total Update Energy Consumption.
Evaluates a machine learning model's ability to detect overshooting tops (OTs) in satellite imagery at a 2 km pixel resolution. It measures how well the model predicts convection/OT presence using physics-informed features derived from visible and infrared channels. Use when the user wants to benchmark on GOES-16 ABI + MRMS Convection Labels, or asks about evaluating this task. Reports hit, correct rejection, false alarm, miss counts.
Evaluates an agent's ability to autonomously plan and execute multi-step GUI automation tasks across various desktop applications. It measures the success rate on both in-distribution tasks from the OSWorld-Verified benchmark and out-of-distribution tasks across six distinct Linux applications. Use when the user wants to benchmark on OSWorld-Verified, OOD GUI Benchmark, or asks about evaluating this task. Reports Success Rate (SR).
Evaluates multimodal agents' ability to perform open-ended, real-world desktop computer tasks across multiple operating systems. It probes GUI grounding, multi-app workflow navigation, and executable action prediction in dynamic, interactive environments. Use when the user wants to benchmark on OSWorld, or asks about evaluating this task. Reports success.
Evaluates deep learning models on classifying osteosarcoma histopathology images into non-tumor, non-viable tumor, viable tumor, and non-viable ratio categories without prior segmentation. Probes the model's ability to capture local texture and global spatial patterns for medical image classification. Use when the user wants to benchmark on TCIA Osteosarcoma, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models' ability to perform online spatio-temporal scene understanding and dynamic, agent-centric reasoning. It tests how well models update spatial and temporal knowledge as they incrementally explore environments, retrieve long-term memory, and infer object relationships across sequential observations. Use when the user wants to benchmark on OST-Bench, or asks about evaluating this task. Reports Overall average score.
Evaluates open semantic mapping models' robustness to dynamic indoor lighting and motion conditions. It measures semantic segmentation accuracy and frequency-weighted IoU, alongside LLM-generated visual question answering accuracy on scene graphs. Use when the user wants to benchmark on ReplicaCAD, HM3D, or asks about evaluating this task. Reports mAcc, f-mIoU.
This benchmark probes a text-to-video model's ability to accurately render and maintain object state transformations (e.g., peeling, slicing) over time. It additionally measures semantic adherence to prompts, scene consistency, and overall perceptual quality to diagnose temporal coherence and physical realism in generated videos. Use when the user wants to benchmark on OSCBench, or asks about evaluating this task. Reports state-change accuracy.
Evaluates subject-driven image generation and manipulation capabilities, specifically testing identity consistency, prompt adherence, and background preservation across single- and multi-subject scenarios. Use when the user wants to benchmark on OSBench, or asks about evaluating this task. Reports Overall (Generation).
This benchmark evaluates optical remote sensing salient object detection models by measuring their ability to accurately segment prominent objects from complex, cluttered backgrounds. It probes structural consistency, boundary precision, and error magnitude across varying object scales and scene complexities. Use when the user wants to benchmark on ORSSD, EORSSD, ORSI-4199, or asks about evaluating this task. Reports maximum F-measure ($F_{\beta}^{max}$).
Evaluates unsupervised time series anomaly detection pipelines across diverse real-world and synthetic datasets. It measures detection accuracy for both point and segment anomalies while tracking computational efficiency and model stability over continuous benchmarking cycles. Use when the user wants to benchmark on OrionBench (NASA, NAB, Yahoo S5, UCR), or asks about evaluating this task. Reports F1 score.
Evaluates models on target-independent stance detection (3-way classification) and argumentative dialogue summarization (overall and stance-specific). It probes the ability to classify conflicting viewpoints in Chinese debates and generate concise, faithful summaries aligned with specific stances. Use when the user wants to benchmark on OrChiD, or asks about evaluating this task. Reports Accuracy, ROUGE-1 F1.
Evaluates Arabic language understanding across seven task clusters, including sentence classification, structured prediction, semantic similarity, NLI, QA, WSD, and topic classification. It probes models' ability to handle diverse Arabic varieties (MSA and dialects) and multiple linguistic levels from tokens to documents. Use when the user wants to benchmark on ORCA, or asks about evaluating this task. Reports ORCA score.
Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.