Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 4,537–4,560 of 23,503 skills
Evaluates the quality of compressed semantic speech representations across automatic speech recognition, speech-to-text translation, and voice conversion tasks. It probes how well adaptive entropy-based token aggregation preserves linguistic and acoustic information under varying compression ratios. Use when the user wants to benchmark on LibriSpeech, CVSS-C, or asks about evaluating this task. Reports WER.
Evaluates the ability of generative latent variable models and autoregressive baselines to model speech audio distributions at varying temporal resolutions. It measures how well models capture intra-frame and inter-frame correlations in audio waveforms by optimizing likelihood objectives. Use when the user wants to benchmark on TIMIT, LibriSpeech, or asks about evaluating this task. Reports bits per frame (bpf).
Evaluates semi-supervised speech enhancement algorithms by measuring how well they recover clean speech from noisy mixtures across varying SNRs and noise types. It probes both perceptual quality and speech intelligibility preservation. Use when the user wants to benchmark on IEEE Speech + Environmental/Industrial Noise Database, or asks about evaluating this task. Reports HASQI.
Evaluates speech enhancement models on their ability to restore audio degraded by multiple distortion types (noise, reverberation, packet loss, clipping, bandwidth limitation, codec artifacts) across varying sampling rates, while preserving speaker identity, phonetic content, and perceptual quality. Use when the user wants to benchmark on DNS 2020 test set, PLC 2024 validation set, VoiceFixer GSR test set, URGENT 2025 non-blind test set, or asks about evaluating this task. Reports DNSMOS.
Evaluates a speech foundation model's ability to generate role-play responses grounded in bottom-up, human-grounded realism. It measures fine-grained aspects such as prosodic dynamics, emotional fidelity, character consistency, and contextual fit based on realistic dialogue and media sources. Use when the user wants to benchmark on DRAME-RoleBench (Realism), or asks about evaluating this task. Reports realism_score.
Evaluates a speech foundation model's ability to generate role-play responses that align with top-down, stereotype-driven character archetypes within specific scene contexts. It measures how well the model captures broad, accessible scoring criteria and general impressions of character consistency without relying on fine-grained prosodic nuance. Use when the user wants to benchmark on DRAME-RoleBench (Archetype), or asks about evaluating this task. Reports archetype_score.
This benchmark evaluates the robustness and cross-domain generalization of speech deepfake detection models across diverse synthetic speech generation techniques, including TTS, voice conversion, neural codecs, and real-world social media leaks. It measures how well models maintain performance when faced with unseen attack types, languages, and distribution shifts. Use when the user wants to benchmark on ASVspoof 2019, ASVspoof 2021, ASVspoof 2024, ADD 2022, ADD 2023, CodecFake, LibriSeVoc, S...
Evaluates a model's ability to distinguish between authentic human speech and synthetically generated or manipulated speech (deepfakes), with a specific focus on robustness against expressive and emotional synthesis attacks. Use when the user wants to benchmark on LibriSpeech, ASVspoof 2019 LA, ASVspoof 2021 LA, ASVspoof 2024, EmoFake, EmoSpoof-TTS, or asks about evaluating this task. Reports EER (Equal Error Rate).
This benchmark probes voice-based and gendered biases in speech continuation models by evaluating how well generated continuations preserve semantic coherence, sentiment, agency, emotional framing, and avoid objectification across different voice qualities (breathy, creaky, end creak) and speaker genders. Use when the user wants to benchmark on SS_set, NOP_set, or asks about evaluating this task. Reports Semantic Coherence.
This benchmark evaluates limited-vocabulary keyword spotting models for on-device speech recognition. It probes a model's ability to correctly identify isolated spoken words from a fixed set of 10 commands, while also handling background silence and unrecognized speech in both aligned and continuous streaming audio contexts. Use when the user wants to benchmark on Speech Commands, or asks about evaluating this task. Reports Top-One Error.
Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying PHP-based e-commerce web workloads. It measures how power consumption scales with session load and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECweb2009, or asks about evaluating this task. Reports watts.
Evaluates multimodal large language models on spectroscopy tasks spanning signal processing, perception, semantic understanding, and molecular generation. It probes the models' ability to align cross-modal data, reason over spectral patterns, and generate accurate chemical structures or spectra. Use when the user wants to benchmark on SpectrumBench, or asks about evaluating this task. Reports accuracy (%).
Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying JVM-based web service workloads. It measures how power consumption scales with workload intensity and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECpower_ssj2008, or asks about evaluating this task. Reports watts.
Evaluates two solver-based algorithms for synthesizing minimal test suites to distinguish between candidate formal specifications (Alloy models). It measures how execution time and test suite size scale with the number of candidate specifications and the domain scope. Use when the user wants to benchmark on Alloy4Fun, or asks about evaluating this task. Reports execution_time.
Evaluates CPU energy efficiency and performance under varying RAPL power caps and core counts using standard SPEC CPU 2017 benchmarks. Probes how power capping influences the trade-off between energy consumption and execution latency across memory-intensive, balanced, and compute-intensive workloads. Use when the user wants to benchmark on SPEC CPU 2017 Speed suite, or asks about evaluating this task. Reports normalized energy usage.
Evaluates whether frontier LLM responses adhere to their published model specifications when faced with generated value tradeoff scenarios. It also measures the consistency of model-based judges in detecting specification violations and identifies specification flaws like contradictions and ambiguities. Use when the user wants to benchmark on Generated Value Tradeoff Scenarios, or asks about evaluating this task. Reports compliance.
Evaluates the quality and latency of streaming text-to-speech models that generate audio incrementally from interleaved text and speech inputs. It probes the model's ability to maintain speech accuracy and naturalness while minimizing first-token latency under streaming constraints. Use when the user wants to benchmark on LJSpeech, LibriSpeech, or asks about evaluating this task. Reports Word Error Rate (WER).
This benchmark evaluates Large Audio-Language Models (LALMs) on their ability to detect, localize, and discriminate speaker inconsistencies in multi-turn dialogues. It specifically probes whether models rely on acoustic cues or are biased toward textual coherence when judging speaker consistency. Use when the user wants to benchmark on SpeakerSleuth, or asks about evaluating this task. Reports detection_accuracy, discrimination_accuracy, localization_f1.
Evaluates a model's ability to convert speech emotion (neutral to angry) while preserving speaker identity across both seen and unseen speakers. It measures spectral and prosody conversion quality objectively, and assesses perceived speech quality, emotion similarity, and speaker similarity subjectively. Use when the user wants to benchmark on English emotional speech corpus, EmoV-DB, JL-Corpus, or asks about evaluating this task. Reports MCD, LSD, PCC.
This benchmark evaluates multimodal large language models on compositional spatial intelligence by testing their ability to reason across 10 atomic spatial capabilities (e.g., counting, localization, spatial relations) combined into 8 complex tasks. It probes scene understanding using 2D and 3D inputs through multiple-choice questions, revealing how models handle hierarchical spatial reasoning and capability integration. Use when the user wants to benchmark on SpaCE-10, or asks about evaluati...
Evaluates clustering algorithms on their ability to track moving and static clusters in collective animal behavior data across space and time. It probes robustness in low-data regimes and the capacity to produce stable, interpretable cluster trajectories without relying on ground-truth labels for hyperparameter tuning. Use when the user wants to benchmark on Cakmak et al. Spatiotemporal Benchmark, or asks about evaluating this task. Reports total AMI.
Evaluates spatiotemporal forecasting models on predicting future frames or climate indices from historical observations. It probes pixel-level reconstruction accuracy, structural similarity, and the model's ability to mitigate error propagation over extended lead times. Use when the user wants to benchmark on Moving MNIST, TrafficBJ, Human 3.6, SEVIR, ICAR-ENSO, or asks about evaluating this task. Reports MSE.
Evaluates the effectiveness of embedding-based spatial keyword retrieval models by measuring how well they rank relevant Points of Interest (POIs) based on combined location and textual query signals. It probes the model's ability to handle spatio-textual relevance without manual weighting of spatial and textual factors. Use when the user wants to benchmark on Beijing, Shanghai, Geo-Glue, or asks about evaluating this task. Reports Recall@k, NDCG@k.
Evaluates the ability of vision-language models to perform multi-step spatial logical reasoning by tracking object dependencies and understanding scene layouts across real-world indoor environments. Use when the user wants to benchmark on SpatiaLQA, or asks about evaluating this task. Reports accuracy.