Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,870
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,1933,216 of 22,870 skills

Vidore EvalA

Evaluates page-level document retrieval on visually rich documents across diverse domains and languages. It probes the model's ability to leverage visual cues, layout, and text within document images without relying on traditional OCR or layout parsing pipelines. Use when the user wants to benchmark on ViDoRe, or asks about evaluating this task. Reports nDCG@5.

researchpython
0
3
Vidmuse Video To Music EvalA

Evaluates a model's ability to generate high-fidelity, semantically aligned music conditioned on video input. It probes audio quality, diversity, and cross-modal alignment using statistical distance metrics, beat alignment scores, and subjective human preference tests. Use when the user wants to benchmark on V2M, AIST++, LORIS, TikTok, or asks about evaluating this task. Reports FAD.

researchpythongo
0
3
Videoscore2 EvalA

This evaluation probes a model's ability to assess generative videos across three key dimensions: visual quality, text-to-video alignment, and physical/common-sense consistency. It measures how well automated scoring models align with human judgments on both in-domain and out-of-domain video benchmarks. Use when the user wants to benchmark on VideoGenReward Bench, T2VQA-DB, MJ-Bench-Video, VideoPhy2-test, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videoscore EvalA

Evaluates how well automatic video quality metrics correlate with human ratings across multiple dimensions such as visual quality, temporal consistency, and text alignment. It also measures pairwise preference accuracy to simulate human choice between generated videos. Use when the user wants to benchmark on VideoFeedback-test, GenAI-Bench, VBench, EvalCrafter, or asks about evaluating this task. Reports Spearman's ρ.

researchpythongo
0
3
Videop2r EvalA

Evaluates large video language models on their ability to perceive visual details and perform multi-step reasoning over video content. It measures how well models decompose video understanding into distinct perception and reasoning stages across multiple benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, VCR, MV, TempCom, VideoMME, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videollama3 EvalA

Evaluates multimodal foundation models on image and video understanding across multiple dimensions, including document and chart text recognition, mathematical reasoning, multi-image comprehension, general knowledge QA, long-form video comprehension, and temporal reasoning. Use when the user wants to benchmark on ChartQA, DocVQA, MathVista, VideoMME, Charades-STA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videohallucer EvalA

This benchmark evaluates large video-language models for intrinsic and extrinsic hallucinations by presenting paired basic and adversarially modified Yes/No questions about video content. It probes whether models can correctly identify factual content while resisting fabricated or unverifiable details, and measures susceptibility to language bias. Use when the user wants to benchmark on VideoHallucer, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Videograin EvalA

Evaluates a video editing model's ability to perform multi-grained edits (class, instance, and part levels) on video-text pairs. It measures semantic alignment, temporal consistency, and pixel-level fidelity of the generated edits. Use when the user wants to benchmark on VideoGrain Evaluation Set, or asks about evaluating this task. Reports CLIP-T.

researchpython
0
3
Videogameqa Bench EvalA

Evaluates vision-language models on video game quality assurance tasks, including glitch detection, temporal reasoning, and bug reporting. It probes the model's ability to process sampled video frames, identify visual anomalies, and generate structured or descriptive reports about game glitches. Use when the user wants to benchmark on VideoGameQA-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Videogamebunny EvalA

Probes vision-language models' ability to understand video game contexts from screenshots, including recognizing actions, characters, UI elements, spatial relationships, and game mechanics. It evaluates how instruction-tuning on game-specific data improves performance compared to larger general-purpose models. Use when the user wants to benchmark on VideoGameBunny Dataset, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Videodpo EvalA

Evaluates text-to-video diffusion models on visual quality and semantic alignment with input prompts. It measures intra-frame fidelity, aesthetic appeal, and inter-frame temporal consistency using automated benchmarks and human-preference predictors. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench.

researchpython
0
3
Videocube EvalA

Evaluates a model's ability to track arbitrary visual instances across complex, unstructured real-world videos without assuming motion continuity or fixed categories. It measures both local search accuracy and global robustness against challenges like occlusion, fast motion, and scene transitions. Use when the user wants to benchmark on VideoCube, or asks about evaluating this task. Reports PRE.

researchpythongo
0
3
Videocraftbench Calvin EvalA

Evaluates a model's ability to learn transferable, long-horizon action dynamics from real-world videos and generate coherent task execution sequences across different environments and robotic setups. Use when the user wants to benchmark on Video-CraftBench, CALVIN, or asks about evaluating this task. Reports Sequential Success Rate (%).

researchpythongo
0
3
Videoconviction EvalA

Evaluates whether LLMs and MLLMs can accurately extract stock tickers, identify explicit investment actions, and quantify human conviction levels from financial influencer videos and transcripts. It probes multimodal reasoning, financial domain understanding, and the ability to filter out noisy or promotional content. Use when the user wants to benchmark on VideoConviction, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Videoaesbench EvalA

Evaluates large multimodal models' ability to perceive and judge video aesthetics across visual form, style, and affectiveness dimensions. It tests performance on diverse video sources (UGC, AIGC, RGC, compression, gaming) using multiple-choice, true/false, and open-ended questions. Use when the user wants to benchmark on VideoAesBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video To Music EvalA

This evaluation protocol assesses a model's ability to generate high-fidelity, diverse instrumental music that is semantically and temporally aligned with a given 10-second video and optional fine-grained text prompt. It probes audio quality, distributional fidelity, generative diversity, and cross-modal alignment using both automated perceptual metrics and human/LLM preference judgments. Use when the user wants to benchmark on ReelBench, LORIS, V2MBench, or asks about evaluating this task. R...

researchpythongo
0
3
Video To C EvalA

Evaluates fine-grained video understanding, spatio-temporal reasoning, and hallucination mitigation in multimodal large language models by testing their ability to locate key visual cues and answer questions across diverse video benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, MVBench, TempCompass, VideoMME, VideoHallucer, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Video To 4d Mesh EvalA

Evaluates a model's ability to generate temporally consistent, animated 3D meshes from input videos. It probes per-frame geometric reconstruction accuracy, overall 4D sequence fidelity, and motion transfer quality while maintaining topology consistency across frames. Use when the user wants to benchmark on Objaverse, Consistent4D, DAVIS, or asks about evaluating this task. Reports CD-3D.

researchpython
0
3
Video Thinking Test EvalA

Evaluates video large language models on their ability to understand complex visual narratives and answer questions correctly. It specifically probes robustness by testing model performance on naturally adversarial or misleading variations of the same video question. Use when the user wants to benchmark on Video Thinking Test, or asks about evaluating this task. Reports Correctness score (accuracy).

researchpythongo
0
3
Video Star EvalA

Evaluates open-vocabulary action recognition by testing a model's ability to generalize to unseen action categories and cross-dataset distributions. It probes fine-grained video understanding and cross-modal reasoning capabilities under base-to-novel and cross-dataset generalization settings. Use when the user wants to benchmark on UCF-101, HMDB-51, Kinetics-400, Kinetics-600, Something-Something V2, or asks about evaluating this task. Reports Top-1 accuracy.

researchpythongo
0
3
Video Reality Test EvalA

This benchmark probes the ability of video-language models and humans to distinguish real ASMR videos from AI-generated ones, evaluating perceptual realism and audio-visual consistency. It also measures how effectively video generation models can deceive video understanding models by producing indistinguishable synthetic content. Use when the user wants to benchmark on Video Reality Test, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Video Qa EvalA

Evaluates video large multimodal models on question-answering tasks across multiple benchmark datasets. Probes the model's ability to understand video content, generate factually accurate long-form responses, and align with language model-derived preferences using direct preference optimization. Use when the user wants to benchmark on MSVD-QA, MSRVTT-QA, TGIF-QA, ActivityNet-QA, VIDAL-QA, WebVid-QA, SSV2-QA, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Video Prediction EvalA

Evaluates the ability of generative models to perform long-horizon open-loop video prediction. It probes how well models maintain temporal consistency, preserve object identities, and adapt to varying scene dynamics across diverse visual domains. Use when the user wants to benchmark on MineRL Navigate, KTH Action, GQN Mazes, Moving MNIST, or asks about evaluating this task. Reports FVD.

researchpython
0
3
Video Phy 2 EvalA

Evaluates the ability of text-to-video models to generate physically plausible content by testing adherence to real-world action-centric physical rules, such as conservation of mass/momentum and gravity. It probes whether models understand fundamental physical commonsense beyond superficial motion or visual aesthetics. Use when the user wants to benchmark on VideoPhy-2, or asks about evaluating this task. Reports joint_performance.

researchpythongo
0
3