Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,865–3,888 of 23,484 skills
Measures the performance, energy efficiency, and latency of the Vega SoC on floating-point near-sensor analytic applications (NSAA) and deep neural network (DNN) inference workloads. Use when the user has predictions and gold and needs to compute Energy Efficiency.
Evaluates video editing reward models and VLM judges on their ability to align with human preferences across three dimensions: instruction following, rendering quality, and edit exclusivity. It tests both global score correlation and local pairwise preference consistency within candidate sets. Use when the user wants to benchmark on VEFX-Bench, or asks about evaluating this task. Reports SRCC.
Evaluates a model's ability to understand and reason about the temporal order of multiple events in long-form videos. It probes whether the model can correctly sequence events, identify relative ordering, and detect pattern anomalies across varying sequence lengths and difficulties. Use when the user wants to benchmark on VECTOR, or asks about evaluating this task. Reports EM (Exact Match).
Evaluates multi-vector visual document retrieval (VDR) models under varying compression ratios. It measures how well pruning and merging strategies maintain retrieval accuracy while reducing storage and computational overhead. Use when the user wants to benchmark on ViDoRe-V1, ViDoRe-V2, JinaVDR-Bench, REAL-MM-RAG, ViDoSeek, MMLongBench-Doc, or asks about evaluating this task. Reports nDCG@5.
Evaluates visually-grounded dialogue systems across five core tasks: multi-modal intent prediction, dialog retrieval (text-to-image and image-to-text), dialog state tracking, and response generation. It provides a unified hierarchical scoring metric to compare cross-task performance and generalization. Use when the user wants to benchmark on VisDial, PhotoChat, MMDialog, Image-Chat, or asks about evaluating this task. Reports VDscore.
This benchmark evaluates video class incremental learning capabilities, testing a model's ability to sequentially learn new action categories while retaining knowledge of previous tasks using limited episodic memory. It specifically probes how well models handle temporal consistency, frame-level memory selection, and classification on both trimmed and untrimmed video data without catastrophic forgetting. Use when the user wants to benchmark on UCF101, Kinetics, ActivityNet-Trim, ActivityNet-U...
Evaluates a model's ability to convert speech from a source speaker to a target speaker within the same language, leveraging limited parallel data alongside a larger non-parallel corpus. The benchmark probes how well systems can disentangle speaker identity from linguistic content when only a small set of aligned sentences is available for training. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
Evaluates a model's ability to convert speech across different languages without parallel data, disentangling speaker characteristics from linguistic content. The benchmark probes cross-lingual generalization and non-parallel training capabilities when target speakers only record in foreign languages. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.
Evaluates non-parallel voice conversion by measuring how effectively a system transfers a source speaker's identity to a target speaker while preserving linguistic content. It probes intra-lingual conversion fidelity and cross-lingual adaptation using subjective human ratings. Use when the user wants to benchmark on VCC2018 (Voice Conversion Challenge 2018), VCTK, or asks about evaluating this task. Reports Quality.
Evaluates voice conversion systems on speech naturalness and target speaker similarity using crowdsourced perceptual tests, and assesses linguistic consistency via automatic speech recognition word error rates. It covers both parallel (Hub) and non-parallel (Spoke) conversion tasks. Use when the user wants to benchmark on VCC2018, or asks about evaluating this task. Reports Naturalness (MOS).
Evaluates voice conversion systems for processing artifacts by repurposing spoofing countermeasures from automatic speaker verification. It measures how easily a detector can distinguish real speech from converted speech, using Equal Error Rate (EER) as a proxy for artifact quality. Use when the user wants to benchmark on VCC'18, or asks about evaluating this task. Reports Equal Error Rate (EER).
Evaluates multimodal mathematical reasoning capabilities of vision-language models, specifically focusing on vision-centric elementary math problems that require explicit visual dependencies across multiple images. It probes spatial, temporal, geometric, logical, and pattern recognition skills to measure how well models integrate cross-modal information for compositional reasoning. Use when the user wants to benchmark on VCBench, or asks about evaluating this task. Reports accuracy.
VCB Bench evaluates audio-grounded large language models on instruction following with speech-level controls, knowledge reasoning, and robustness under real-world acoustic perturbations. It probes how well models understand and generate spoken responses in Chinese and English using authentic human speech rather than synthetic data. Use when the user wants to benchmark on VCB Bench, or asks about evaluating this task. Reports 1-5 scale score.
Evaluates one-shot voice conversion quality by measuring how naturally the converted speech sounds and how closely it matches the target speaker's voice compared to human baselines. Use when the user wants to benchmark on VCTK, LibriTTS, or asks about evaluating this task. Reports MOS (Naturalness & Similarity).
Evaluates the factual accuracy and overall quality of video captions in a reference-free setting. It measures how well a model's predicted quality scores and explanations align with human judgments across diverse video and image domains. Use when the user has predictions and gold and needs to compute Kendall's correlation ($\tau_b$).
Evaluates multimodal large language models' ability to follow instructions containing vision-dependent constraints, such as spatial, stylistic, and structural requirements. It isolates the contribution of visual input to instruction adherence and assesses generalization on standard visual reasoning tasks. Use when the user wants to benchmark on VC-IFEval, MM-IFEval, IFEval, or asks about evaluating this task. Reports instruction-following accuracy.
Systematic video reasoning capabilities grounded in five cognitive faculties: perception, transformation, spatiality, abstraction, and knowledge. It probes spatiotemporal reasoning, mental manipulation, and rule-based problem solving on video sequences. Use when the user wants to benchmark on VBVR-Dataset, or asks about evaluating this task. Reports rule-based scorer.
Evaluates video language models by disentangling question types into LLM-Answerable, Semantic, Temporal, and Others. It isolates true temporal and spatial understanding from language priors and static visual cues by computing accuracy exclusively on the Semantic and Temporal subsets. Use when the user wants to benchmark on LongVideoBench, Egoschema, NextQA, VideoMME, MLVU, LVBench, PerceptionTest, or asks about evaluating this task. Reports VBenchComp score.
Evaluates text-to-video generation quality and semantic alignment across dimensions like human action, scene composition, object consistency, and aesthetic quality. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Overall.
Evaluates the visual quality, semantic alignment, temporal consistency, and aesthetic fidelity of text-to-video generation models. It probes the model's ability to produce coherent, high-fidelity videos that match textual prompts across multiple perceptual and technical dimensions. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench Total Score.
Evaluates the quality and trustworthiness of text-to-video and image-to-video generative models across 16 fine-grained dimensions, including spatial consistency, temporal dynamics, and subject identity. It measures how well automated scores align with human preferences and compares frame-wise generation capabilities against text-to-image baselines. Use when the user wants to benchmark on VBench++, or asks about evaluating this task. Reports VBench score.
Evaluates abstractive headline generation across 15 Indic languages and English. It probes cross-lingual transfer, script normalization effects, and the impact of language-family-specific pretraining on low-resource generation. Use when the user wants to benchmark on Varta, or asks about evaluating this task. Reports ROUGE-L.
Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).