Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,265–3,288 of 22,874 skills
Evaluates voice conversion systems on speech naturalness and target speaker similarity using crowdsourced perceptual tests, and assesses linguistic consistency via automatic speech recognition word error rates. It covers both parallel (Hub) and non-parallel (Spoke) conversion tasks. Use when the user wants to benchmark on VCC2018, or asks about evaluating this task. Reports Naturalness (MOS).
Evaluates voice conversion systems for processing artifacts by repurposing spoofing countermeasures from automatic speaker verification. It measures how easily a detector can distinguish real speech from converted speech, using Equal Error Rate (EER) as a proxy for artifact quality. Use when the user wants to benchmark on VCC'18, or asks about evaluating this task. Reports Equal Error Rate (EER).
Evaluates multimodal mathematical reasoning capabilities of vision-language models, specifically focusing on vision-centric elementary math problems that require explicit visual dependencies across multiple images. It probes spatial, temporal, geometric, logical, and pattern recognition skills to measure how well models integrate cross-modal information for compositional reasoning. Use when the user wants to benchmark on VCBench, or asks about evaluating this task. Reports accuracy.
VCB Bench evaluates audio-grounded large language models on instruction following with speech-level controls, knowledge reasoning, and robustness under real-world acoustic perturbations. It probes how well models understand and generate spoken responses in Chinese and English using authentic human speech rather than synthetic data. Use when the user wants to benchmark on VCB Bench, or asks about evaluating this task. Reports 1-5 scale score.
Evaluates one-shot voice conversion quality by measuring how naturally the converted speech sounds and how closely it matches the target speaker's voice compared to human baselines. Use when the user wants to benchmark on VCTK, LibriTTS, or asks about evaluating this task. Reports MOS (Naturalness & Similarity).
Evaluates the factual accuracy and overall quality of video captions in a reference-free setting. It measures how well a model's predicted quality scores and explanations align with human judgments across diverse video and image domains. Use when the user has predictions and gold and needs to compute Kendall's correlation ($\tau_b$).
Evaluates multimodal large language models' ability to follow instructions containing vision-dependent constraints, such as spatial, stylistic, and structural requirements. It isolates the contribution of visual input to instruction adherence and assesses generalization on standard visual reasoning tasks. Use when the user wants to benchmark on VC-IFEval, MM-IFEval, IFEval, or asks about evaluating this task. Reports instruction-following accuracy.
Systematic video reasoning capabilities grounded in five cognitive faculties: perception, transformation, spatiality, abstraction, and knowledge. It probes spatiotemporal reasoning, mental manipulation, and rule-based problem solving on video sequences. Use when the user wants to benchmark on VBVR-Dataset, or asks about evaluating this task. Reports rule-based scorer.
Evaluates video language models by disentangling question types into LLM-Answerable, Semantic, Temporal, and Others. It isolates true temporal and spatial understanding from language priors and static visual cues by computing accuracy exclusively on the Semantic and Temporal subsets. Use when the user wants to benchmark on LongVideoBench, Egoschema, NextQA, VideoMME, MLVU, LVBench, PerceptionTest, or asks about evaluating this task. Reports VBenchComp score.
Evaluates text-to-video generation quality and semantic alignment across dimensions like human action, scene composition, object consistency, and aesthetic quality. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Overall.
Evaluates the visual quality, semantic alignment, temporal consistency, and aesthetic fidelity of text-to-video generation models. It probes the model's ability to produce coherent, high-fidelity videos that match textual prompts across multiple perceptual and technical dimensions. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Total Score.
Evaluates the quality and trustworthiness of text-to-video and image-to-video generative models across 16 fine-grained dimensions, including spatial consistency, temporal dynamics, and subject identity. It measures how well automated scores align with human preferences and compares frame-wise generation capabilities against text-to-image baselines. Use when the user wants to benchmark on VBench++, or asks about evaluating this task. Reports VBench score.
Evaluates abstractive headline generation across 15 Indic languages and English. It probes cross-lingual transfer, script normalization effects, and the impact of language-family-specific pretraining on low-resource generation. Use when the user wants to benchmark on Varta, or asks about evaluating this task. Reports ROUGE-L.
Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).
Evaluates the ability of video-language models to detect subtle, rapid, and contextually nuanced anomalies in both real-world surveillance footage and high-fidelity AI-generated videos. It probes fine-grained temporal reasoning, visual grounding, and robustness to synthetic artifacts through a multiple-choice question-answering format. Use when the user wants to benchmark on VANE-Bench, or asks about evaluating this task. Reports MC-Video QA accuracy.
Evaluates whether multimodal large language models (MLLMs) can maintain consistent culture-conditioned value judgments when response options are replaced with minimally contrastive visual proxies. It probes cross-modal prediction stability and the ability to ground textual value tendencies in subtle visual contrasts. Use when the user wants to benchmark on ValueGround, or asks about evaluating this task. Reports accuracy.
Evaluates the faithfulness and semantic plausibility of model-agnostic XAI techniques by generating soft counterfactuals via token-level perturbations. It measures whether perturbations actually change model predictions and whether the generated explanations align with the true causal impact of those changes. Use when the user has predictions and gold and needs to compute Validitysoft, Csoft.
This protocol evaluates the perceptual fidelity and cross-domain generalization capability of the VALERIE22 synthetic urban dataset by training a semantic segmentation model on it and testing on real-world automotive datasets. It specifically probes how dataset diversity (unique 3D assets) and training scale affect downstream perception performance. Use when the user wants to benchmark on VALERIE22, Cityscapes, A2D2, BDD100K, India Driving Dataset, Mapillary Vistas, or asks about evaluating t...
Evaluates multimodal large language models' ability to perform extractive and abstractive spatiotemporal reasoning on egocentric videos. It probes long-horizon memory, object tracking, spatial orientation, and metric distance estimation under both multiple-choice and free-form generation settings. Use when the user wants to benchmark on VAEX-Bench, or asks about evaluating this task. Reports Accuracy.
Evaluates the effectiveness of Variational Autoencoder (VAE)-derived latent space features for malware classification using traditional machine learning models. It probes robustness to data partitioning, random seed initialization, and computational efficiency without hyperparameter tuning. Use when the user wants to benchmark on EMBER, BODMAS, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict continuous emotional dimensions (Valence, Arousal, Dominance) from text. Specifically probes the model's capacity to capture affective polarization signals in parliamentary discourse. Use when the user wants to benchmark on Knesset VAD Annotation, or asks about evaluating this task. Reports Pearson correlation.
Evaluates a model's ability to detect anomalous events in surveillance videos and anticipate their occurrence in future frames. It specifically probes scene-dependent anomaly recognition and multi-step temporal anticipation. Use when the user wants to benchmark on ShanghaiTech, CUHK Avenue, IITB Corridor, NWPU Campus, ShanghaiTech-sd, or asks about evaluating this task. Reports AUC (%).
Evaluates the quality, cross-modal consistency, synchronization, and spatial audio rendering of text-to-audio-video and image-to-audio-video generation models. It probes physical plausibility, emotional expressiveness, and stereo separation across seven real-world sound categories. Use when the user wants to benchmark on VABench, or asks about evaluating this task. Reports Audio-Visual Align.
This evaluation protocol assesses the utility of the Vaani dataset for fine-tuning automatic speech recognition (ASR) and spoken language identification (LID) models across diverse Indian languages and regions. It measures performance gains from fine-tuning on Vaani's transcribed audio and images against established benchmarks, highlighting regional dialectal variations and low-resource language capabilities. Use when the user wants to benchmark on Vaani, FLEURS, Kathbath, or asks about evalu...