Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,217–3,240 of 22,874 skills
This benchmark probes the ability of video-language models and humans to distinguish real ASMR videos from AI-generated ones, evaluating perceptual realism and audio-visual consistency. It also measures how effectively video generation models can deceive video understanding models by producing indistinguishable synthetic content. Use when the user wants to benchmark on Video Reality Test, or asks about evaluating this task. Reports accuracy.
Evaluates video large multimodal models on question-answering tasks across multiple benchmark datasets. Probes the model's ability to understand video content, generate factually accurate long-form responses, and align with language model-derived preferences using direct preference optimization. Use when the user wants to benchmark on MSVD-QA, MSRVTT-QA, TGIF-QA, ActivityNet-QA, VIDAL-QA, WebVid-QA, SSV2-QA, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of generative models to perform long-horizon open-loop video prediction. It probes how well models maintain temporal consistency, preserve object identities, and adapt to varying scene dynamics across diverse visual domains. Use when the user wants to benchmark on MineRL Navigate, KTH Action, GQN Mazes, Moving MNIST, or asks about evaluating this task. Reports FVD.
Evaluates the ability of text-to-video models to generate physically plausible content by testing adherence to real-world action-centric physical rules, such as conservation of mass/momentum and gravity. It probes whether models understand fundamental physical commonsense beyond superficial motion or visual aesthetics. Use when the user wants to benchmark on VideoPhy-2, or asks about evaluating this task. Reports joint_performance.
Evaluates a model's ability to generate spatially and temporally consistent video content outside the original frame boundaries (video outpainting), while preserving source structure and visual realism. Use when the user wants to benchmark on DAVIS 2017, YouTube-VOS, or asks about evaluating this task. Reports FVD.
This protocol audits video understanding benchmarks to measure genuine spatio-temporal reasoning versus shortcut reliance. It filters out samples solvable without video context and evaluates models under diagnostic conditions (e.g., blind, audio-only, center-frame) to quantify performance degradation and dependency on actual video content. Use when the user wants to benchmark on EgoSchema, ImplicitQA, VSI-Bench, TVBench, VCR-Bench, RTV-Bench, Video-Holmes, MINERVA, MMR-V, VideoMME, MVBench, L...
Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on Video-MME, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to segment moving and unknown objects in autonomous driving videos without relying on a closed set of known classes. It probes open-set and motion-based instance segmentation capabilities under varying data distributions and synthetic scenarios. Use when the user wants to benchmark on Cityscapes-VPS, KITTI-MOTS, Carla, or asks about evaluating this task. Reports CAQ, CA-IoU.
Evaluates fine-grained audiovisual captioning quality, attribute-level instruction following, and downstream reasoning capabilities like QA and temporal grounding. Use when the user wants to benchmark on video-SALMONN-2, UGC-VideoCap, VDC, VidCapBench-AE, Daily-Omni, World-Sense, Charades-STA, or asks about evaluating this task. Reports accuracy.
Evaluates video generation models across two core dimensions: video-condition alignment (how well the generated video matches the text prompt in terms of objects, actions, colors, scenes, and overall consistency) and video quality (technical fidelity, aesthetics, temporal consistency, and motion quality). Use when the user wants to benchmark on Video-Bench, or asks about evaluating this task. Reports Video-text Consistency.
Evaluates long-video understanding and agentic retrieval capabilities by testing models on temporal grounding, summarization, reasoning, and event causality across ultra-long video streams. Use when the user wants to benchmark on LVBench, VideoMME-Long, Ava-100, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to recognize and classify human actions in video clips by predicting action categories from sampled frames. It probes temporal dynamics modeling and spatial feature extraction capabilities in video understanding tasks. Use when the user wants to benchmark on Kinetics-400, Something-Something-V2, Epic-Kitchens-100, or asks about evaluating this task. Reports Top-1 accuracy.
Egocentric video understanding for embodied AI, probing capabilities in video question-answering, hierarchical task planning, visual grounding, and reward modeling. It evaluates how well multimodal models comprehend first-person, action-oriented video contexts required for robotic interaction. Use when the user wants to benchmark on VidEgoThink, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to assess danger levels in videos by identifying risk elements, understanding context, and assigning severity scores. It probes multimodal perception and risk reasoning capabilities. Use when the user wants to benchmark on ViDAS, or asks about evaluating this task. Reports MSE.
Probes image-level logical anomaly detection under vision-induced distractions such as background changes, blur, and low light. It tests whether models can identify violations of logical constraints (e.g., quantity, length, type, placement) by reasoning over textual descriptions rather than relying on brittle low-level visual features. Use when the user wants to benchmark on VID-AD, or asks about evaluating this task. Reports AUROC.
This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.
Evaluates LLMs on fault-targeted test generation and fault-targeted program repair. It probes discriminative fault detection, fault hypothesis generation, and the ability to debug subtle semantic bugs under diagnostic guidance. Use when the user wants to benchmark on VIBEPASS, or asks about evaluating this task. Reports D_{IO}.
Probes multimodal reasoning and visual understanding on real-world images. Specifically designed with a 'hard' subset of prompts that are unsolvable by current frontier models to measure genuine performance gaps and contamination-free generalization. Use when the user wants to benchmark on Vibe-Eval, or asks about evaluating this task. Reports Vibe-Eval Score.
Evaluates the ability of Vision-Language Models to avoid generating unfaithful or non-existent visual details in both closed-set discriminative QA and open-ended generative captioning. It also measures how well an auxiliary probing framework can detect these hallucinations via attention dynamics and mitigate them at inference time without retraining. Use when the user wants to benchmark on POPE, AMBER, M-HalDetect, COCO-Caption, or asks about evaluating this task. Reports AUPRC.
Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.
Evaluates multimodal models' ability to detect harmful content in images and videos across ten specific harmful categories and a general unharmful class. It probes binary classification robustness against dataset imbalance and multi-class reasoning capabilities under varying prompt conditions. Use when the user wants to benchmark on VHD11K, SMID, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates language-based image segmentation by requiring models to ground natural language phrases into precise image regions. It probes a model's ability to handle long-tail categories, attributes, relationships, and varying object sizes in open-vocabulary settings. Use when the user wants to benchmark on VGPhraseCut, or asks about evaluating this task. Reports mean-IoU.
Evaluates audio-visual foundation and embedding models on multi-label video classification, probing their ability to recognize sound and visual events across different input modalities. It specifically measures modality alignment, unimodal versus multimodal performance, and susceptibility to distraction from irrelevant background audio or static visuals. Use when the user wants to benchmark on VGGSounder, or asks about evaluating this task. Reports F1-score.
Evaluates zero-shot language-queried audio source separation on human actions, sound-emitting objects, and human-object interactions. The benchmark tests isolation of a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on VGGSound, or asks about evaluating this task. Reports SDRi.