Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,073–6,096 of 23,905 skills
Evaluates a model's ability to detect out-of-distribution (OOD) inputs in 3D point cloud semantic segmentation. It probes domain shift robustness (indoor vs outdoor scenes) and sensor failure simulation (missing color channels) by measuring uncertainty-based OOD scores against in-distribution data. Use when the user wants to benchmark on Semantic3D, S3DIS, Semantic3D (no color), or asks about evaluating this task. Reports AUROC.
Out-of-domain molecular property prediction and Bayesian optimization for molecular design. Tests transferability of learned representations to novel tasks. Use when the user wants to benchmark on Out-of-domain molecular design tasks, or asks about evaluating this task. Reports Top performing molecule property.
Evaluates the ability of uncertainty estimation methods to distinguish in-distribution clinical histopathology samples from out-of-distribution samples and to provide well-calibrated confidence scores for selective prediction. It probes how different methods maintain predictive accuracy and calibration under varying degrees of data shift. Use when the user wants to benchmark on CIFAR-10, Hospital 4, Hospital 5, Cancer type, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to distinguish between in-distribution (ID) samples and out-of-distribution (OOD) samples. It measures how well the model's decision boundaries separate known classes from unknown data distributions using various synthetic and natural OOD benchmarks. Use when the user wants to benchmark on SVHN, CIFAR-10, CIFAR-100, TinyImageNet, TinyImageNet-crop, TinyImageNet-resize, LSUN-crop, LSUN-resize, iSUN, or asks about evaluating this task. Reports TNR@TPR95.
Evaluates a model's ability to distinguish in-distribution chest X-rays from out-of-distribution medical images (e.g., knee, hand, or general radiographs) while maintaining classification accuracy on chest diseases. Use when the user wants to benchmark on CXR14, IRMA, MURA, Bone Age, or asks about evaluating this task. Reports AUC.
Evaluates the ability of out-of-distribution (OOD) detection methods to identify distribution shifts in 3D medical image segmentation. It measures how well models distinguish in-distribution scans from clinically anomalous or shifted OOD scans, highlighting the limitations of deep learning-based detectors compared to simpler intensity-based baselines. Use when the user wants to benchmark on 3D CT datasets, 3D MRI datasets, or asks about evaluating this task. Reports FPR at 95% TPR (FPR95).
Evaluates large language models' ability to perform ontology subsumption inference by framing it as a binary natural language inference task. The model must predict whether a hypothesis concept subsumes a premise concept based on verbalized OWL axioms. Use when the user wants to benchmark on biMNLI, Schema.org (Atomic SI), DOID (Atomic SI), FoodOn (Atomic SI), GO (Atomic SI), FoodOn (Complex SI), GO (Complex SI), or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the accuracy of mapping OpenAlex paper topics to terms across 13 scientific ontologies. It probes a method's ability to perform semantic and lexical alignment between informal topic labels and formal domain-specific vocabularies. Use when the user wants to benchmark on Ontology Alignment Gold Standard, or asks about evaluating this task. Reports F1.
Evaluates a sports-domain language model's generation capability on sports-specific tasks and its zero-shot commonsense reasoning performance on general benchmarks. Use when the user wants to benchmark on OnlySports Benchmark, HellaSwag, PIQA, ARC-challenge, ARC-easy, or asks about evaluating this task. Reports OS-acc.
Evaluates creative problem-solving and associative reasoning by testing whether models can correctly group words and identify connections, specifically probing susceptibility to cognitive fixation effects when misleading red herring clues are present. Use when the user wants to benchmark on Only Connect Wall (OCW), or asks about evaluating this task. Reports grouping_evaluation.
Evaluates a model's ability to perform real-time, streaming video understanding and long-horizon reasoning. It probes low-latency perception, temporal alignment, evidence-aligned response timing, and causal reasoning across short to hour-long video sequences. Use when the user wants to benchmark on StreamingBench, OVOBench, RTVBench, OVBench, VideoMME, MLVU, LongVideoBench, LVBench, or asks about evaluating this task. Reports QA accuracy.
Evaluates online matching algorithms for maximizing individual and group fairness among offline agents, as well as weighted matching performance, under dynamic arrival constraints. Use when the user wants to benchmark on Chicago ride-hailing dataset, Network Data Repository (socfb-Caltech36, socfb-Reed98, econ-because, econ-mbeaflw), Synthetic bipartite graphs, or asks about evaluating this task. Reports CR1.
This protocol evaluates an online feature interaction detection method integrated into click-through rate (CTR) prediction models. It measures predictive accuracy on streaming ad click data using chronological splits to simulate real-time recommendation scenarios. Use when the user wants to benchmark on Avazu, Criteo, Taobao, or asks about evaluating this task. Reports AUC, logloss.
Evaluates a unified multimodal reasoning model's ability to perform visual understanding tasks across both static images and videos. It probes capabilities in question answering, captioning, spatial and temporal grounding, object tracking, and segmentation. Use when the user wants to benchmark on MMMU, MathVista, MathVerse, MMBench, MMStar, ScienceQA, AI2D, MMT-Bench, VideoMMMU, MMVU, VideoMME, VideoHolmes, LongVideoBench, LongVideo-Reason, VideoMathQA, MMSci-Caption, MMT-Caption, VideoMMLU-C...
Evaluates a generative recommendation model's ability to perform sequential next-item prediction and in-text reasoning for user preference alignment. It probes the model's capacity to generate interpretable reasoning paths alongside item recommendations, and measures ranking accuracy on standard recommendation benchmarks. Use when the user wants to benchmark on Amazon Beauty, Amazon Toys, Amazon Sports, or asks about evaluating this task. Reports R@K / NDCG@K.
Measures broader image generation capabilities across five axes: alignment, text rendering, reasoning, style, and diversity. Use when the user wants to benchmark on OneIG-Bench, or asks about evaluating this task. Reports Overall.
Evaluates a model's ability to extract structured information from chart images, including textual OCR accuracy and precise numerical value parsing. It tests the model's capacity to convert visual chart elements into a standardized Python-dict representation, handling both annotated and unannotated charts across multiple languages and rendering styles. Use when the user wants to benchmark on ChartQA-SE, PlotQA-SE, ChartX-SE, ChartY-en, ChartY-zh, or asks about evaluating this task. Reports SC...
This evaluation probes a robot's ability to generalize a single kinesthetic demonstration to novel object poses and orientations using unseen object pose estimation for trajectory transfer. It measures how robustly different pose estimation methods enable successful completion of everyday manipulation tasks in real-world settings. Use when the user wants to benchmark on Custom 10-task real-world manipulation set, or asks about evaluating this task. Reports success rate (%).
This benchmark evaluates a system's ability to perform one-shot information extraction from document images. It measures how accurately the model can extract specific entity values (e.g., dates, amounts, names) from unseen test documents after being shown only a single training example, with optional supplementary documents for refinement. Use when the user wants to benchmark on Doctor's Bills, Patent (Ghega), or asks about evaluating this task. Reports extraction accuracy.
Evaluates the trade-offs between inference speed, energy consumption, latency, and generation quality of various LLM architectures and quantization schemes running on mobile hardware. It specifically measures how model size, sparsity (MoE), and compression formats impact physical battery drain and user-perceived responsiveness. Use when the user wants to benchmark on Summarization Task (On-Device Profiling), or asks about evaluating this task. Reports Energy per Token (Joules).
Evaluates an LLM agent's ability to generate high-quality, multi-objective ad keywords from product descriptions. It probes lexical and semantic alignment with product info, real-world campaign performance (clicks, conversions, cost), and the model's capacity for self-reflective refinement over multiple generation rounds. Use when the user wants to benchmark on OKG Benchmark Dataset, or asks about evaluating this task. Reports ROUGE-1.
Evaluates 3D geometric foundation models on monocular and video depth estimation, and tests camera-controlled video generation models on their ability to follow camera trajectories while maintaining video quality across diverse, dynamic environments. Use when the user wants to benchmark on OmniWorld-Game, or asks about evaluating this task. Reports FVD.
This evaluation probes the viewpoint invariance and robustness of vision-language pre-training models. It measures how well models maintain classification accuracy on clean data, common out-of-distribution shifts, and specifically challenging viewpoint-variant images compared to standard baselines. Use when the user wants to benchmark on ImageNet-1K, ImageNet-V+, ImageNet-V, OOD-CV, MIRO, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates multimodal large language models' ability to jointly reason across visual and audio modalities in long-duration videos. It probes capabilities like cross-modal alignment, temporal dependency modeling, and understanding of low-semantic acoustic cues such as music and ambient sounds. Use when the user wants to benchmark on OmniVideoBench, or asks about evaluating this task. Reports accuracy.