Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,913–3,936 of 23,482 skills
This benchmark probes the ability of large language models to generate rigorous, step-by-step mathematical proofs for high-school olympiad-level problems. It evaluates logical coherence, justification of assumptions, and adherence to formal proof standards rather than just numerical correctness. Use when the user wants to benchmark on 2025 USA Math Olympiad, or asks about evaluating this task. Reports proof_points.
Evaluates the ability of deep learning architectures (SSMs, Transformers, RNNs) to forecast hourly electricity load across major US power grids. It probes how well models capture temporal patterns, handle varying prediction horizons, and integrate exogenous weather covariates for accurate grid-scale forecasting. Use when the user wants to benchmark on US ISO Hourly Load Data (EIA-930), or asks about evaluating this task. Reports MSE (%).
Evaluates cross-domain speech recognition and enhancement robustness by training downstream models on generatively simulated target-domain data. Probes the ability of ASR and SE systems to generalize to unseen acoustic conditions, channel mismatches, and compound noise-channel distortions. Use when the user wants to benchmark on Hakka Across Taiwan (HAT), Taiwanese Across Taiwan (TAT), VoiceBank-DEMAND (VBD), HAT-ESC, or asks about evaluating this task. Reports CER.
Evaluates end-to-end speech-to-speech dialogue models across three core dimensions: understanding, reasoning, and oral conversation. It probes multilingual proficiency, multi-turn dialogue handling, and the ability to generate paralinguistic and emotional cues in audio responses. Use when the user wants to benchmark on URO-Bench, or asks about evaluating this task. Reports Task Accomplish Score.
Evaluates how well uncertainty estimates for latent representations predict the correctness of the representation itself, specifically whether the nearest neighbor in the embedding space belongs to the same class. It tests the transferability and scalability of uncertainty quantification methods across different backbones and unseen datasets. Use when the user wants to benchmark on ImageNet-1k, or asks about evaluating this task. Reports R-AUROC.
Evaluates the fidelity of automated urban scene reconstruction from city-tour videos and the generalization capability of reinforcement learning navigation policies trained in these simulations. It measures how well generated scenes match real-world semantics and layouts, and how effectively policies transfer to unseen simulated and real-world environments. Use when the user wants to benchmark on KITTI-360, CraftBench, AutoBench, or asks about evaluating this task. Reports success_rate.
Evaluates the capability of segmentation models to detect and delineate individual tree crowns and canopy coverage from high-resolution aerial imagery across diverse urban and tropical environments. Use when the user wants to benchmark on Zurich Municipal Tree Inventory & Swisstopo Imagery, WeRobotics Open AI Challenge (Tonga), or asks about evaluating this task. Reports Recall.
Evaluates the utility of a semi-procedurally generated synthetic driving dataset (UrbanSyn) for unsupervised domain adaptation (UDA) in semantic segmentation. It probes whether combining multiple synthetic sources reduces the domain gap and improves pixel-level classification accuracy on real-world urban driving benchmarks. Use when the user wants to benchmark on UrbanSyn, GTAV, Synscapes, Cityscapes, BDD100K, Mapillary Vistas, or asks about evaluating this task. Reports self-labeling accuracy.
Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark on Berlin Urban Simulation, or asks about evaluating this task. Reports Average Travel Time (s).
This benchmark evaluates Machine Reading Comprehension (MRC) capabilities in Urdu by testing a model's ability to extract correct answer spans from context paragraphs in response to questions. It probes span prediction accuracy, handling of multiple valid answers, and performance across different question types and named entities. Use when the user wants to benchmark on UQuAD1.0, or asks about evaluating this task. Reports F1.
Evaluates recommendation accuracy and counterfactual fairness of LLM-based recommendation models. It measures ranking performance using Hit@k metrics and assesses bias by calculating the AUC for predicting sensitive user attributes from recommendations. Use when the user wants to benchmark on MovieLens-1M, Insurance, or asks about evaluating this task. Reports Hit@1.
Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of language models to perform unsupervised relation extraction by predicting relation labels or tokens from contextual text. It probes factual grounding and context-constrained generation capabilities across varying relation types and corpus sources. Use when the user wants to benchmark on T-REx, Google-RE, ZSRE, TACRED, or asks about evaluating this task. Reports F1.
Evaluates the ability of image descriptors to distinguish near-duplicate image pairs from non-duplicates under extreme specificity constraints, simulating large-scale forensic or fraud detection scenarios. Use when the user wants to benchmark on MFND (Mir-Flickr Near-Duplicate), CLAIMS, Holidays, California-ND, or asks about evaluating this task. Reports sensitivity at false positive rate (FPR).
This benchmark evaluates a model's ability to detect unsupervised lexical semantic change across diachronic corpus pairs. It probes two capabilities: binary classification of whether a word's sense has been gained or lost, and ranking the intensity of semantic change relative to a gold standard. Use when the user wants to benchmark on SemEval-2020 Task 1, or asks about evaluating this task. Reports accuracy, Spearman’s rank-order correlation coefficient.
Evaluates whether language models exhibit gender bias when processing sentence pairs that have been filtered to remove explicit gendered language and stereotypical co-occurrences. It measures the model's ability to generate gender-neutral completions and checks for systematic preference toward male or female pronouns in stereotype-free contexts. Use when the user wants to benchmark on USE-5, USE-10, USE-20, WB (Winobias), WG (Winogender), or asks about evaluating this task. Reports US fairnes...
Evaluates a model's ability to recognize emotions in speech from speakers it has never encountered during training. It probes cross-speaker generalization and robustness to acoustic variability across multiple languages and recording conditions. Use when the user wants to benchmark on CREMA-D, IEMOCAP, RAVDESS, EmoDB, CaFE, BhavVani, or asks about evaluating this task. Reports WF1.
Evaluates a model's ability to estimate the 6D pose (rotation and translation) of novel, unseen 3D objects in real-world scenes without retraining, using only their mesh models and partial RGBD inputs. It specifically probes robustness to pose ambiguity, partial observability, and real-world noise. Use when the user wants to benchmark on GraspNet-1Billion, YCB-Video, or asks about evaluating this task. Reports IADD.
Evaluates the effectiveness of image safety classifiers in detecting various unsafe content categories across real-world and AI-generated images. It also probes classifier robustness to distribution shifts caused by artistic representations and grid layouts in AI-generated content. Use when the user wants to benchmark on UnsafeBench, or asks about evaluating this task. Reports F1-Score.
Evaluates embodied AI agents' capabilities in complex, photo-realistic 3D open-world environments. Specifically probes visual navigation on unstructured terrain, active visual tracking across diverse scenes, and social tracking under dynamic distractions, varying morphologies, and different control frequencies. Use when the user wants to benchmark on UnrealZoo, or asks about evaluating this task. Reports Success Rate (SR).
Assesses multimodal language models' ability to resolve lexical ambiguity in puns using visual context. It probes visual-textual alignment, multimodal literacy, and the capacity to disambiguate or reconstruct ambiguous text when provided with explanatory or disambiguating images. Use when the user wants to benchmark on UNPIE, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates the ability of recommender systems to efficiently remove specific user interactions or sensitive items (unlearning) while preserving recommendation utility. It probes real-world operational constraints, including handling sequential small-batch deletion requests, domain-specific triggers, and low-latency execution across collaborative filtering, session-based, and next-basket recommendation tasks. Use when the user wants to benchmark on TaFeng, Dunnhumby, Instacart, RSC15, DIGI, NOW...
Evaluates a model's ability to recognize human actions from heterogeneous skeleton data with varying joint counts and topologies. It probes cross-domain generalization, zero-shot/few-shot transfer, and robustness to structural discrepancies between sensing modalities. Use when the user wants to benchmark on NTU-60, HumanML3D, NW-UCLA, NTU-120, or asks about evaluating this task. Reports Accuracy.
Evaluates multilingual named entity recognition (NER) capabilities across 22 languages and 30 datasets, probing both in-language performance and cross-lingual transfer. It also benchmarks large language models as annotators against human inter-annotator agreement to assess guideline adherence and annotation quality. Use when the user wants to benchmark on UNER v2, or asks about evaluating this task. Reports micro F1.