Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,503
skills in category
980
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,417–4,440 of 23,503 skills

Stylebench EvalA

Evaluates speech language models' ability to control four paralinguistic dimensions (emotion, speed, volume, pitch) across multi-turn dialogues. It probes whether models can follow gradational style-intensity instructions while preserving fixed semantic content and maintaining coherent intensity trajectories across turns. Use when the user wants to benchmark on StyleBench, or asks about evaluating this task. Reports dimension-specific metrics.

researchpythongo
0
3
Style4rec EvalA

Evaluates a model's ability to perform sequential product recommendation in an e-commerce setting by predicting the next item in a user session. It specifically probes how well the model leverages historical interaction sequences, visual style embeddings, and shopping cart data to rank relevant products. Use when the user wants to benchmark on Style4Rec E-commerce Dataset, or asks about evaluating this task. Reports HR@5.

researchpythonperformance
0
3
Style Transfer EvalA

Evaluates unsupervised style transfer models on their ability to transform text between two styles (e.g., Shakespeare vs. modern English, formal vs. informal) while preserving semantic meaning and maintaining linguistic quality. Use when the user wants to benchmark on Shakespeare author imitation dataset (Xu et al., 2012), Formality transfer dataset (Rao and Tetrault, 2018), or asks about evaluating this task. Reports J(A,S,F).

researchpython
0
3
Stvg EvalA

Evaluates a model's ability to localize a specific object or event in a video based on a natural language query. It measures both temporal localization (identifying the correct start and end timestamps) and spatial localization (predicting accurate bounding box trajectories across the video frames). Use when the user wants to benchmark on VidSTG, HCSTVG-v1&v2, or asks about evaluating this task. Reports m_vIoU.

researchpythongo
0
3
Structured3d Layout EvalA

Evaluates a model's ability to predict architectural elements (walls, doors, windows) and room layouts within indoor 3D scenes. It tests the model's capacity for structured scene understanding and spatial reasoning by comparing predicted layouts against ground-truth annotations. Use when the user wants to benchmark on Structured3D, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Structured Prompting EvalA

Evaluates how different prompting strategies (baseline, zero-shot, chain-of-thought, and automated optimizers) affect the accuracy, ranking stability, and variance of language model performance across multiple knowledge and reasoning benchmarks. Use when the user wants to benchmark on MMLU-Pro, GSM8K, MedCalc-Bench, GPQA, HeadQA, MedBullets, Medec, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Structured Output Benchmark EvalA

Evaluates large language models' ability to extract structured information from multi-modal sources (text, images, audio) into valid JSON formats, isolating schema compliance from value accuracy. Use when the user wants to benchmark on Multi-Source Structured Output Benchmark, or asks about evaluating this task. Reports correct_value_extraction.

researchpythongo
0
3
Structured Data Qa EvalA

Evaluates the ability of LLMs to generate executable structured queries (SQL, SPARQL) from natural language questions over tables, knowledge graphs, and temporal knowledge graphs. It specifically probes the model's capacity for iterative self-correction when initial query generation fails, measuring both raw generation accuracy and error-resilience during multi-round inference. Use when the user wants to benchmark on WikiSQL, WTQ, MetaQA, WebQSP, CronQuestions, TabFact, or asks about evaluati...

researchpythongo
0
3
Struct Bench EvalA

This benchmark evaluates the quality of differentially private synthetic structured text generation by measuring how well synthetic datasets preserve the syntactic structure, semantic dependencies, and attribute distributions of real data, alongside downstream task utility. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports CFG Pass Rate (CFG-PR).

researchpythongo
0
3
Streetfighter EvalA

Tests an LLM agent's real-time decision-making and combat strategy in a video game environment, prioritizing low latency while maintaining competitive win rates. The benchmark probes the model's ability to make timely character actions under a hard frame-rate limit. Use when the user wants to benchmark on StreetFighter, or asks about evaluating this task. Reports ELO Score.

researchpythongit
0
3
Streamuni EvalA

Evaluates real-time speech translation systems on latency and translation quality across multiple language pairs. It probes the model's ability to dynamically decide when to generate and truncate translations while processing streaming audio inputs. Use when the user wants to benchmark on MuST-C English→German, MuST-C English→Spanish, CoVoST2 English→Chinese, CoVoST2 French→English, or asks about evaluating this task. Reports SacreBLEU, COMET.

researchpythonapi
0
3
Streammark EvalA

This benchmark evaluates the imperceptibility, robustness to benign audio transformations, and semi-fragility to malicious deepfake manipulations of a deep learning-based audio watermarking system. It specifically probes whether a watermark can survive standard compression and cropping while being deliberately destroyed by semantic-altering AI conversions like voice cloning or speech editing. Use when the user wants to benchmark on LibriSpeech (train_clean100), Test Set A, Test Set B, or asks...

researchpythongo
0
3
Streamingbench EvalA

Evaluates multimodal large language models' ability to understand real-time streaming video, integrate visual and audio information, maintain contextual continuity across sequential questions, and proactively output information at specific timestamps. Use when the user wants to benchmark on StreamingBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Streaming 3d Reconstruction EvalA

Evaluates a model's ability to perform streaming camera pose estimation and 3D reconstruction over long video sequences. It probes long-range geometric consistency, drift resistance, and reconstruction fidelity across diverse indoor and outdoor environments. Use when the user wants to benchmark on Oxford Spires, ETH3D, 7-Scenes, Tanks and Temples, NRGBD, or asks about evaluating this task. Reports ATE, F1.

researchpythongo
0
3
Streamflow Hydrology EvalA

Evaluates hydrological forecasting models on predicting streamflow across 319 US basins. It tests the model's ability to capture long-term temporal dependencies and intermediate hydrological states (soil water, snowpack) under varying data segmentation, training sizes, and noise conditions. Use when the user wants to benchmark on 319 US Basins, or asks about evaluating this task. Reports RMSE.

researchpythonperformance
0
3
Streambench EvalA

Evaluates real-time streaming video understanding and multi-turn dialogue capabilities. It measures semantic correctness, factual accuracy, dialogue coherence, and system latency across diverse video types and query formats. Use when the user wants to benchmark on STREAMBENCH, or asks about evaluating this task. Reports Acc..

researchpythongo
0
3
StreamA

Evaluates the spatial realism and temporal flow consistency of videos generated by video generative models. It decomposes video features into spatial and temporal components using embedding spaces and Fourier transforms to provide length-agnostic, bounded scores that independently measure visual fidelity and motion naturalness. Use when the user has predictions and gold and needs to compute STREAM-T.

researchpythongo
0
3
Stream Omni EvalA

Evaluates multimodal capabilities across visual understanding, speech interaction, and vision-grounded speech tasks. It measures accuracy on standard VQA and knowledge-grounded QA benchmarks, and uses LLM-based scoring for open-ended spoken interactions. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, SEED-Bench, LLaVA-Bench-in-the-Wild, MM-Vet, Llama Questions, Web Questions, SpokenVisIT, or asks about evaluating this task. Reports acc...

researchpythongo
0
3
Stormnet Bias EvalA

Evaluates a spatio-temporal graph neural network's ability to predict and correct systematic biases in storm surge water level forecasts. It probes the model's capacity to leverage spatial dependencies among coastal gauge stations and temporal patterns to improve long-horizon (up to 72h) hydrodynamic predictions. Use when the user wants to benchmark on Gulf Coast Gauge Network (NOAA/TCOON), or asks about evaluating this task. Reports RMSE.

researchpythonperformance
0
3
Stormcast Forecast EvalA

Evaluates the ability of a generative diffusion model to emulate kilometer-scale atmospheric convection and precipitation forecasting. It probes spatial-temporal forecast skill, multi-variable physical consistency (e.g., updrafts, cold pools), and spectral fidelity over 1-6 hour lead times. Use when the user wants to benchmark on ERA5, HRRR, MRMS, or asks about evaluating this task. Reports FSS.

researchpython
0
3
Stochastic Ackley EvalA

Evaluates the ability of uncertainty-aware deep neural networks to approximate a highly irregular, multi-extremum function and quantify predictive uncertainty. It specifically probes how well the models handle in-distribution versus out-of-distribution parameter regimes. Use when the user wants to benchmark on Stochastic Ackley Function, or asks about evaluating this task. Reports Relative Error (RE).

researchpythongo
0
3
Stgformer EvalA

This evaluation protocol tests a model's ability to forecast future traffic flow across large-scale urban road networks using historical spatiotemporal sensor data. It probes the model's capacity to capture complex spatial dependencies and temporal dynamics while maintaining computational efficiency on real-world traffic benchmarks. Use when the user wants to benchmark on LargeST, PEMS-series, or asks about evaluating this task. Reports MAE.

researchpythonperformance
0
3
Stg4traffic EvalA

Evaluates the multi-step spatio-temporal forecasting capability of graph neural networks on urban traffic speed and flow prediction tasks. It measures how well models capture spatial dependencies and temporal dynamics across varying prediction horizons (3, 6, and 12 steps). Use when the user wants to benchmark on METR-LA, PEMS-BAY, PEMSD4, PEMSD8, or asks about evaluating this task. Reports MAE.

researchpythonperformance
0
3
Stf Extraction EvalA

Evaluates a model's ability to recover coherent source time functions (STFs) from scattered, noisy seismic wavefields without relying on traditional deconvolution or labeled seismograms. Use when the user wants to benchmark on Synthetic Scattering Simulation, or asks about evaluating this task. Reports maximum normalized cross-correlation (MNCC).

researchpythongo
0
3