Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,417–4,440 of 23,503 skills
Evaluates speech language models' ability to control four paralinguistic dimensions (emotion, speed, volume, pitch) across multi-turn dialogues. It probes whether models can follow gradational style-intensity instructions while preserving fixed semantic content and maintaining coherent intensity trajectories across turns. Use when the user wants to benchmark on StyleBench, or asks about evaluating this task. Reports dimension-specific metrics.
Evaluates a model's ability to perform sequential product recommendation in an e-commerce setting by predicting the next item in a user session. It specifically probes how well the model leverages historical interaction sequences, visual style embeddings, and shopping cart data to rank relevant products. Use when the user wants to benchmark on Style4Rec E-commerce Dataset, or asks about evaluating this task. Reports HR@5.
Evaluates unsupervised style transfer models on their ability to transform text between two styles (e.g., Shakespeare vs. modern English, formal vs. informal) while preserving semantic meaning and maintaining linguistic quality. Use when the user wants to benchmark on Shakespeare author imitation dataset (Xu et al., 2012), Formality transfer dataset (Rao and Tetrault, 2018), or asks about evaluating this task. Reports J(A,S,F).
Evaluates a model's ability to localize a specific object or event in a video based on a natural language query. It measures both temporal localization (identifying the correct start and end timestamps) and spatial localization (predicting accurate bounding box trajectories across the video frames). Use when the user wants to benchmark on VidSTG, HCSTVG-v1&v2, or asks about evaluating this task. Reports m_vIoU.
Evaluates a model's ability to predict architectural elements (walls, doors, windows) and room layouts within indoor 3D scenes. It tests the model's capacity for structured scene understanding and spatial reasoning by comparing predicted layouts against ground-truth annotations. Use when the user wants to benchmark on Structured3D, or asks about evaluating this task. Reports F1.
Evaluates how different prompting strategies (baseline, zero-shot, chain-of-thought, and automated optimizers) affect the accuracy, ranking stability, and variance of language model performance across multiple knowledge and reasoning benchmarks. Use when the user wants to benchmark on MMLU-Pro, GSM8K, MedCalc-Bench, GPQA, HeadQA, MedBullets, Medec, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to extract structured information from multi-modal sources (text, images, audio) into valid JSON formats, isolating schema compliance from value accuracy. Use when the user wants to benchmark on Multi-Source Structured Output Benchmark, or asks about evaluating this task. Reports correct_value_extraction.
Evaluates the ability of LLMs to generate executable structured queries (SQL, SPARQL) from natural language questions over tables, knowledge graphs, and temporal knowledge graphs. It specifically probes the model's capacity for iterative self-correction when initial query generation fails, measuring both raw generation accuracy and error-resilience during multi-round inference. Use when the user wants to benchmark on WikiSQL, WTQ, MetaQA, WebQSP, CronQuestions, TabFact, or asks about evaluati...
This benchmark evaluates the quality of differentially private synthetic structured text generation by measuring how well synthetic datasets preserve the syntactic structure, semantic dependencies, and attribute distributions of real data, alongside downstream task utility. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports CFG Pass Rate (CFG-PR).
Tests an LLM agent's real-time decision-making and combat strategy in a video game environment, prioritizing low latency while maintaining competitive win rates. The benchmark probes the model's ability to make timely character actions under a hard frame-rate limit. Use when the user wants to benchmark on StreetFighter, or asks about evaluating this task. Reports ELO Score.
Evaluates real-time speech translation systems on latency and translation quality across multiple language pairs. It probes the model's ability to dynamically decide when to generate and truncate translations while processing streaming audio inputs. Use when the user wants to benchmark on MuST-C English→German, MuST-C English→Spanish, CoVoST2 English→Chinese, CoVoST2 French→English, or asks about evaluating this task. Reports SacreBLEU, COMET.
This benchmark evaluates the imperceptibility, robustness to benign audio transformations, and semi-fragility to malicious deepfake manipulations of a deep learning-based audio watermarking system. It specifically probes whether a watermark can survive standard compression and cropping while being deliberately destroyed by semantic-altering AI conversions like voice cloning or speech editing. Use when the user wants to benchmark on LibriSpeech (train_clean100), Test Set A, Test Set B, or asks...
Evaluates multimodal large language models' ability to understand real-time streaming video, integrate visual and audio information, maintain contextual continuity across sequential questions, and proactively output information at specific timestamps. Use when the user wants to benchmark on StreamingBench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform streaming camera pose estimation and 3D reconstruction over long video sequences. It probes long-range geometric consistency, drift resistance, and reconstruction fidelity across diverse indoor and outdoor environments. Use when the user wants to benchmark on Oxford Spires, ETH3D, 7-Scenes, Tanks and Temples, NRGBD, or asks about evaluating this task. Reports ATE, F1.
Evaluates hydrological forecasting models on predicting streamflow across 319 US basins. It tests the model's ability to capture long-term temporal dependencies and intermediate hydrological states (soil water, snowpack) under varying data segmentation, training sizes, and noise conditions. Use when the user wants to benchmark on 319 US Basins, or asks about evaluating this task. Reports RMSE.
Evaluates real-time streaming video understanding and multi-turn dialogue capabilities. It measures semantic correctness, factual accuracy, dialogue coherence, and system latency across diverse video types and query formats. Use when the user wants to benchmark on STREAMBENCH, or asks about evaluating this task. Reports Acc..
Evaluates the spatial realism and temporal flow consistency of videos generated by video generative models. It decomposes video features into spatial and temporal components using embedding spaces and Fourier transforms to provide length-agnostic, bounded scores that independently measure visual fidelity and motion naturalness. Use when the user has predictions and gold and needs to compute STREAM-T.
Evaluates multimodal capabilities across visual understanding, speech interaction, and vision-grounded speech tasks. It measures accuracy on standard VQA and knowledge-grounded QA benchmarks, and uses LLM-based scoring for open-ended spoken interactions. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, SEED-Bench, LLaVA-Bench-in-the-Wild, MM-Vet, Llama Questions, Web Questions, SpokenVisIT, or asks about evaluating this task. Reports acc...
Evaluates a spatio-temporal graph neural network's ability to predict and correct systematic biases in storm surge water level forecasts. It probes the model's capacity to leverage spatial dependencies among coastal gauge stations and temporal patterns to improve long-horizon (up to 72h) hydrodynamic predictions. Use when the user wants to benchmark on Gulf Coast Gauge Network (NOAA/TCOON), or asks about evaluating this task. Reports RMSE.
Evaluates the ability of a generative diffusion model to emulate kilometer-scale atmospheric convection and precipitation forecasting. It probes spatial-temporal forecast skill, multi-variable physical consistency (e.g., updrafts, cold pools), and spectral fidelity over 1-6 hour lead times. Use when the user wants to benchmark on ERA5, HRRR, MRMS, or asks about evaluating this task. Reports FSS.
Evaluates the ability of uncertainty-aware deep neural networks to approximate a highly irregular, multi-extremum function and quantify predictive uncertainty. It specifically probes how well the models handle in-distribution versus out-of-distribution parameter regimes. Use when the user wants to benchmark on Stochastic Ackley Function, or asks about evaluating this task. Reports Relative Error (RE).
This evaluation protocol tests a model's ability to forecast future traffic flow across large-scale urban road networks using historical spatiotemporal sensor data. It probes the model's capacity to capture complex spatial dependencies and temporal dynamics while maintaining computational efficiency on real-world traffic benchmarks. Use when the user wants to benchmark on LargeST, PEMS-series, or asks about evaluating this task. Reports MAE.
Evaluates the multi-step spatio-temporal forecasting capability of graph neural networks on urban traffic speed and flow prediction tasks. It measures how well models capture spatial dependencies and temporal dynamics across varying prediction horizons (3, 6, and 12 steps). Use when the user wants to benchmark on METR-LA, PEMS-BAY, PEMSD4, PEMSD8, or asks about evaluating this task. Reports MAE.
Evaluates a model's ability to recover coherent source time functions (STFs) from scattered, noisy seismic wavefields without relying on traditional deconvolution or labeled seismograms. Use when the user wants to benchmark on Synthetic Scattering Simulation, or asks about evaluating this task. Reports maximum normalized cross-correlation (MNCC).