Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,049–6,072 of 23,917 skills
Evaluates the capability of models to detect machine-generated text and attribute it to specific authors or models across diverse scenarios, including mixed human-machine text, out-of-domain content, and outputs from unseen LLMs. Use when the user wants to benchmark on OpenTuringBench, or asks about evaluating this task. Reports F1-score.
This benchmark probes Theory-of-Mind (ToM) reasoning in LLMs by testing their ability to infer psychological mental states (e.g., beliefs, attitudes, intentions) and track physical object locations across naturally generated narratives. It specifically evaluates first- and second-order ToM capabilities under varying narrative lengths and question types. Use when the user wants to benchmark on OpenToM, or asks about evaluating this task. Reports macro-averaged F1 score.
Evaluates deep learning models on patch-based and patient-level multiclass classification of brain tumor histology images. It probes the ability of CNNs and vision transformers to distinguish between different tumor types and normal tissue using intraoperative stimulated Raman histology data. Use when the user wants to benchmark on OpenSRH, or asks about evaluating this task. Reports top-1 accuracy.
Evaluates open-world audio source separation by measuring how well a model disentangles multiple audio sources from a mixed input. It probes the model's ability to generalize to seen and unseen audio classes and handle complex natural mixtures without manual intervention. Use when the user wants to benchmark on MUSIC, VGGSound, AudioCaps, or asks about evaluating this task. Reports SDR.
Evaluates the capability of web search agents to perform multi-step navigation, complex deep research planning, and precise information retrieval across English and Chinese web environments. It probes the model's ability to synthesize information from noisy, long-horizon browsing trajectories and extract exact answers or reliable summaries. Use when the user wants to benchmark on BrowseComp, BrowseComp-ZH, xbench-DeepSearch, WideSearch, or asks about evaluating this task. Reports accuracy / F...
This benchmark evaluates the ability of vision models to detect and localize diffusion-generated images in an open-world setting. It probes cross-domain generalization across multiple diffusion architectures (SD1.5, SD2.1, SDXL, SD3, Flux.1) and tests robustness against common image degradations like Gaussian blur and JPEG compression. Use when the user wants to benchmark on OpenSDID, or asks about evaluating this task. Reports F1.
Evaluates 3D vision models' ability to perform open-vocabulary scene understanding by querying for fine-grained object attributes (e.g., material, affordance, synonym) rather than standard object classes. It measures both 3D instance segmentation and 3D semantic segmentation capabilities on attribute-based queries across eight linguistic aspects. Use when the user wants to benchmark on OpenScan, or asks about evaluating this task. Reports AP, mIoU.
Evaluates Subject-to-Video (S2V) generation models on their ability to maintain subject identity consistency, produce natural temporal dynamics, and align with text prompts. It covers open-domain, human-specific, and single-subject scenarios to expose common failure modes like copy-paste artifacts and fidelity degradation. Use when the user wants to benchmark on OpenS2V-Eval, or asks about evaluating this task. Reports NexusScore, NaturalScore, GmeScore.
This benchmark evaluates the recommendation capability of LLM-based systems on sequential and straightforward recommendation tasks. It probes how well models leverage user interaction histories and different item indexing strategies to predict relevant items across multiple public datasets. Use when the user wants to benchmark on Movielens-1M, Amazon Beauty, LastFM, or asks about evaluating this task. Reports HR@k, NDCG@k.
Evaluates named entity recognition (NER) capabilities across 52 languages and 36 distinct corpora. It probes cross-lingual generalization, robustness to varying entity type ontologies, and the ability of both encoder-based models and LLMs to handle multilingual text with diverse annotation guidelines. Use when the user wants to benchmark on OpenNER 1.0, or asks about evaluating this task. Reports micro-averaged mention-level F1.
Evaluates the energy efficiency and performance of OpenMP loop transformations (tiling, unrolling) and parallel constructs across different compilers and workloads. Use when the user wants to benchmark on Matrix Multiplication, 2D Stencil, Barcelona OpenMP Task Suite (BOTS), NAS Parallel Benchmarks, PARSEC benchmark, or asks about evaluating this task. Reports Energy (J).
Evaluates machine learning classifiers on a curated collection of standardized classification tasks. It probes the reproducibility and comparability of algorithm performance across diverse datasets under consistent, machine-readable evaluation protocols. Use when the user wants to benchmark on OpenML-CC18, or asks about evaluating this task. Reports accuracy_score.
This benchmark evaluates the performance and cross-framework compatibility of deep learning algorithms for medical image analysis across classification, segmentation, localization, and detection tasks. It specifically probes how model accuracy and inference efficiency vary when implementations are ported between PyTorch and MindSpore on heterogeneous hardware (NVIDIA GPUs vs. Huawei Ascend NPUs). Use when the user wants to benchmark on OpenMedIA Benchmark Suite, or asks about evaluating this ...
Evaluates the quality of open-ended interleaved image-text generation by comparing model outputs against human-annotated references. It probes multimodal coherence, visual fidelity, and text-image alignment through pairwise battles and automated judge agreement. Use when the user wants to benchmark on OpenING, or asks about evaluating this task. Reports agreement.
Evaluates Open Information Extraction systems on their ability to accurately identify and extract relational triples from natural language sentences. It measures precision, recall, and F1 using multiple reference-matching protocols, alongside throughput speed and confidence-threshold robustness (AUC). Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports F1.
Evaluates the performance and efficiency of state-of-the-art neural OpenIE models and training datasets across multiple standard benchmarks. It probes how model properties like N-ary relation support and inferred relation extraction capability align with benchmark characteristics and downstream task requirements. Use when the user wants to benchmark on OIE2016, WiRE57, ReOIE2016, CaRB, LSOIE, or asks about evaluating this task. Reports F1 score.
This evaluation probes a model's ability to perform Open Information Extraction (OpenIE), which involves identifying and extracting relational triples (subject, predicate, object) from unstructured text without relying on a predefined ontology or schema. It measures how well systems capture complete, correct, and minimal information spans across diverse domains like news and encyclopedias. Use when the user wants to benchmark on OIE2016, CaRB, or asks about evaluating this task. Reports F1.
Evaluates deep learning models for seismic full-waveform inversion (FWI) by predicting subsurface velocity models from seismic wavefield data. It probes the model's ability to generalize across varying geological complexities and out-of-distribution scenarios using parameter-efficient fine-tuning. Use when the user wants to benchmark on OpenFWI, or asks about evaluating this task. Reports SSIM.
Evaluates the predictive accuracy and computational efficiency of a pruned deep learning model for seismic full waveform inversion. It measures how well the model reconstructs subsurface velocity maps and quantifies inference latency and resource consumption on edge hardware. Use when the user wants to benchmark on OpenFWI, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to predict open-vocabulary functional 3D scene graphs from posed RGB-D images. It specifically probes the detection of objects and interactive elements, as well as the inference of their functional relationships (e.g., switch controls light) in real-world indoor spaces. Use when the user wants to benchmark on SceneFun3D, FunGraph3D, or asks about evaluating this task. Reports Recall@K.
Evaluates language models' ability to make probabilistic forecasts on open-ended, future-uncertain questions derived from global news. It probes both prediction accuracy and calibration, testing whether models can generalize forecasting skills across diverse sources and time horizons without leaking future information. Use when the user wants to benchmark on OpenForesight Test Set, FutureX, SimpleQA, MMLU-Pro, GPQA-Diamond, or asks about evaluating this task. Reports Brier Score.
Binary classification capability for detecting AI-generated images versus real photographs. It probes a model's ability to generalize across diverse generative models (diffusion, transformer-based) and real-world social media distributions. Use when the user wants to benchmark on OpenFake, or asks about evaluating this task. Reports F1 Score.
Probes legal reasoning capabilities in U.S. bankruptcy exemption law, specifically testing multi-step inference, robustness to distractors and obfuscation, and scalability across asset counts and temporal complexity. Use when the user wants to benchmark on OpenExempt, or asks about evaluating this task. Reports macro-averaged F1.
Probes a model's ability to perform multimodal event grounding by integrating visual content with textual context. It evaluates three interdependent capabilities: generating event-aware image captions, retrieving relevant news articles from images, and retrieving images from narrative event descriptions. Use when the user wants to benchmark on OpenEvents V1, or asks about evaluating this task. Reports CLIPScore, mAP.