Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,241–3,264 of 22,874 skills
Evaluates a model's ability to perform continual audio-visual classification across sequential tasks without catastrophic forgetting. It measures how well the model retains performance on previously learned categories while learning new ones, across audio, visual, and cross-modal fusion settings. Use when the user wants to benchmark on VGGSound-Instruments, VGGSound-100, VGG-Sound Source, or asks about evaluating this task. Reports Average accuracy.
Evaluates a vision-language model's ability to understand graphical user interfaces (GUIs) and answer user questions based on visual content. It specifically probes the model's capacity to avoid hallucinations by grounding responses in actual GUI elements rather than relying solely on textual priors. Use when the user wants to benchmark on GUI Comprehension Bench, or asks about evaluating this task. Reports GPT evaluation score.
Evaluates text-to-video generation models across three dimensions: aesthetic quality, aesthetic tagging, and generation fidelity. It measures how well automated neural assessors align with human judgments and ranks models based on normalized scores across multiple visual and formal dimensions. Use when the user wants to benchmark on VGA-Bench, or asks about evaluating this task. Reports five-class accuracy.
Evaluates the visual reasoning and grounding capabilities of Large Vision-Language Models (LVLMs) by measuring the quality of their step-by-step rationales, the accuracy of their final answers, and the alignment between the generated reasoning and the prediction. Use when the user wants to benchmark on VG-CoT, or asks about evaluating this task. Reports Rationale Quality (RQ), Answer Accuracy (AA), Reasoning-Answer Alignment (RAA).
This benchmark evaluates a model's ability to perform zero-shot voice imitation by disentangling linguistic content, speaker timbre, and vocal style (accent/emotion). It probes the model's capacity to generate high-intelligibility speech that accurately transfers the target speaker's identity and stylistic attributes from a reference clip without task-specific fine-tuning. Use when the user wants to benchmark on Vevo Evaluation Set (AB, CV, ACCENT, EMOTION), or asks about evaluating this task...
VEU-Bench evaluates a model's ability to understand video editing by probing 19 fine-grained tasks across 10 dimensions (e.g., shot size, cut types, transitions) and three cognitive stages: recognition, reasoning, and judging. It tests whether models can identify editing components, infer their functions, and judge their effects in video content. Use when the user wants to benchmark on VEU-Bench, or asks about evaluating this task. Reports Scoreall.
Evaluates a unified multimodal model's ability to generate synchronized audio and video from text, phonemes, and reference media. It probes zero-shot voice cloning fidelity, lip-sync accuracy, acoustic quality, and cross-modal temporal alignment. Use when the user wants to benchmark on VerseBench, or asks about evaluating this task. Reports WER.
Evaluates general visual reasoning capabilities across a diverse set of 30 benchmarks spanning six task categories, including chart/OCR, STEM, spatial/action, knowledge/recognition, grounding, and captioning/instruction following. Use when the user wants to benchmark on VeroEval, or asks about evaluating this task. Reports overall averages.
Evaluates multimodal misinformation detection models on real-world and synthetic image-caption pairs, specifically probing their ability to distinguish truthful content from out-of-context (OOC) and miscaptioned (MC) misinformation while measuring susceptibility to unimodal bias. Use when the user wants to benchmark on VERITE, COSMOS, VMU-Twitter, or asks about evaluating this task. Reports accuracy.
This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.
Evaluates the factual correctness of long-form LLM-generated responses by decomposing them into atomic facts, detecting and refining incomplete or missing information, and verifying each fact against external web evidence. Use when the user wants to benchmark on Long-form LLM responses, or asks about evaluating this task. Reports Supported/Contradicted/Undecided classification accuracy.
Evaluates the reasoning capabilities of voice and multimodal models under real-time streaming constraints, quantifying the performance gap between text and voice modalities on tasks with well-defined ground truth. Use when the user wants to benchmark on VERA, or asks about evaluating this task. Reports Accuracy.
This benchmark evaluates GUI grounding capabilities across a hierarchical taxonomy of basic (element, visual, spatial) and advanced (functional, reasoning, refusal) tasks. It probes a model's ability to accurately locate UI elements in screenshots and handle complex, domain-specific, or unanswerable instructions across web, mobile, and desktop platforms. Use when the user wants to benchmark on VenusBench-GD, or asks about evaluating this task. Reports accuracy.
Evaluates a U-Net model's ability to predict velocity fold numbers and produce dealiased radar velocity fields from folded inputs. It measures both classification accuracy for fold detection and reconstruction fidelity via velocity error metrics. Use when the user wants to benchmark on WSR-88D Level-II/III Radar Data, or asks about evaluating this task. Reports velocity RMSE.
Measures the performance, energy efficiency, and latency of the Vega SoC on floating-point near-sensor analytic applications (NSAA) and deep neural network (DNN) inference workloads. Use when the user has predictions and gold and needs to compute Energy Efficiency.
Evaluates video editing reward models and VLM judges on their ability to align with human preferences across three dimensions: instruction following, rendering quality, and edit exclusivity. It tests both global score correlation and local pairwise preference consistency within candidate sets. Use when the user wants to benchmark on VEFX-Bench, or asks about evaluating this task. Reports SRCC.
Evaluates a model's ability to understand and reason about the temporal order of multiple events in long-form videos. It probes whether the model can correctly sequence events, identify relative ordering, and detect pattern anomalies across varying sequence lengths and difficulties. Use when the user wants to benchmark on VECTOR, or asks about evaluating this task. Reports EM (Exact Match).
Evaluates multi-vector visual document retrieval (VDR) models under varying compression ratios. It measures how well pruning and merging strategies maintain retrieval accuracy while reducing storage and computational overhead. Use when the user wants to benchmark on ViDoRe-V1, ViDoRe-V2, JinaVDR-Bench, REAL-MM-RAG, ViDoSeek, MMLongBench-Doc, or asks about evaluating this task. Reports nDCG@5.
Evaluates visually-grounded dialogue systems across five core tasks: multi-modal intent prediction, dialog retrieval (text-to-image and image-to-text), dialog state tracking, and response generation. It provides a unified hierarchical scoring metric to compare cross-task performance and generalization. Use when the user wants to benchmark on VisDial, PhotoChat, MMDialog, Image-Chat, or asks about evaluating this task. Reports VDscore.
This benchmark evaluates video class incremental learning capabilities, testing a model's ability to sequentially learn new action categories while retaining knowledge of previous tasks using limited episodic memory. It specifically probes how well models handle temporal consistency, frame-level memory selection, and classification on both trimmed and untrimmed video data without catastrophic forgetting. Use when the user wants to benchmark on UCF101, Kinetics, ActivityNet-Trim, ActivityNet-U...
Evaluates a model's ability to convert speech from a source speaker to a target speaker within the same language, leveraging limited parallel data alongside a larger non-parallel corpus. The benchmark probes how well systems can disentangle speaker identity from linguistic content when only a small set of aligned sentences is available for training. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
Evaluates a model's ability to convert speech across different languages without parallel data, disentangling speaker characteristics from linguistic content. The benchmark probes cross-lingual generalization and non-parallel training capabilities when target speakers only record in foreign languages. Use when the user wants to benchmark on EMIME, or asks about evaluating this task. Reports MOS (Mean Opinion Score).
Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.
Evaluates non-parallel voice conversion by measuring how effectively a system transfers a source speaker's identity to a target speaker while preserving linguistic content. It probes intra-lingual conversion fidelity and cross-lingual adaptation using subjective human ratings. Use when the user wants to benchmark on VCC2018 (Voice Conversion Challenge 2018), VCTK, or asks about evaluating this task. Reports Quality.