Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,793–3,816 of 23,483 skills
This benchmark evaluates automatic speech recognition (ASR) models on Vietnamese medical audio containing embedded English terminology. It specifically probes the model's ability to accurately transcribe both the matrix language and code-switched segments, measuring overall transcription quality alongside specialized metrics for code-switched and non-code-switched spans. Use when the user wants to benchmark on ViMedCSS, or asks about evaluating this task. Reports WER.
This evaluation probes a model's ability to generalise to novel robotic manipulation tasks by testing robustness to instruction variations and increased task difficulty. It specifically measures compositional generalisation capabilities across four systematicity levels, ranging from object pose sensitivity to entirely novel objects and tasks. Use when the user wants to benchmark on VIMABench, or asks about evaluating this task. Reports compositional generalisation capabilities.
Evaluates video-language continual learning by testing a model's ability to retain episodic memories across streaming, long-duration videos without catastrophic forgetting. It probes cross-modal inference and temporal localization across three non-classification tasks: moment queries, natural language queries, and visual queries. Use when the user wants to benchmark on ViLCo-Bench, or asks about evaluating this task. Reports Average Recall@k (IoU=m).
Evaluates vision-language models' ability to solve multi-hop visual reasoning tasks by measuring how accurately their final predicted answers match the ground truth. It specifically probes the model's capacity for structured reasoning and answer extraction in complex domains like geometry, science, and visual question answering. Use when the user wants to benchmark on MAVIS-Geometry, A-OKVQA, GeoQA170K, CLEVR-Math, ScienceQA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to extract robust visual features from single-channel mammography images for binary classification of malignant versus benign breast lumps. It tests cross-dataset generalization and the effectiveness of multimodal contrastive pretraining on pathological classification tasks. Use when the user wants to benchmark on MVKL, CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect hate and offensive text spans within Vietnamese social media comments. It probes sequence tagging capabilities, specifically requiring precise boundary identification of offensive content in noisy, informal text. Use when the user wants to benchmark on ViHOS, or asks about evaluating this task. Reports macro-average F1-score.
This evaluation probes a model's ability to perform scene change detection conditioned on natural language prompts. It measures how well the model distinguishes relevant semantic changes from nuisance variations across diverse domains (street-view, satellite, indoor) and handles viewpoint misalignments. Use when the user wants to benchmark on CSeg, PSCD, SYSU-CD, VL-CMU-CD, or asks about evaluating this task. Reports IoU.
Identifies and categorizes abusive content spans within long-form Vietnamese narrative texts. It probes a model's ability to perform sequence labeling for both span detection and fine-grained abuse classification across six distinct categories. Use when the user wants to benchmark on Vietnamese Narrative Abusive Span Dataset, or asks about evaluating this task. Reports Strict F-score.
Evaluates abstractive summarization capabilities on real-world and simulated medical conversations in Vietnamese, testing both human-transcribed and ASR-generated noisy transcripts. Use when the user wants to benchmark on VietMed-Sum, or asks about evaluating this task. Reports ROUGE.
Evaluates the ability of NER models to identify and classify medically defined entity spans in Vietnamese spoken text. It specifically probes robustness to ASR-generated noise and compares monolingual vs. multilingual, encoder vs. seq2seq architectures. Use when the user wants to benchmark on VietMed-NER, or asks about evaluating this task. Reports micro F1 score.
Evaluates automatic speech recognition (ASR) performance on Vietnamese medical domain audio. It measures how well models transcribe speech containing medical terminology and regional accents, assessing cross-domain transfer capabilities. Use when the user wants to benchmark on VietMed, or asks about evaluating this task. Reports WER.
This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.
This benchmark evaluates end-to-end Retrieval Augmented Generation (RAG) systems on visually rich, real-world documents across multiple professional domains. It probes a model's ability to retrieve relevant pages, generate accurate answers to complex open-ended and multi-hop queries, and precisely ground those answers with bounding boxes in multimodal content. Use when the user wants to benchmark on ViDoRe V3, or asks about evaluating this task. Reports F1 score (Dice coefficient).
Evaluates page-level document retrieval on visually rich documents across diverse domains and languages. It probes the model's ability to leverage visual cues, layout, and text within document images without relying on traditional OCR or layout parsing pipelines. Use when the user wants to benchmark on ViDoRe, or asks about evaluating this task. Reports nDCG@5.
Evaluates a model's ability to generate high-fidelity, semantically aligned music conditioned on video input. It probes audio quality, diversity, and cross-modal alignment using statistical distance metrics, beat alignment scores, and subjective human preference tests. Use when the user wants to benchmark on V2M, AIST++, LORIS, TikTok, or asks about evaluating this task. Reports FAD.
This evaluation probes a model's ability to assess generative videos across three key dimensions: visual quality, text-to-video alignment, and physical/common-sense consistency. It measures how well automated scoring models align with human judgments on both in-domain and out-of-domain video benchmarks. Use when the user wants to benchmark on VideoGenReward Bench, T2VQA-DB, MJ-Bench-Video, VideoPhy2-test, or asks about evaluating this task. Reports accuracy.
Evaluates how well automatic video quality metrics correlate with human ratings across multiple dimensions such as visual quality, temporal consistency, and text alignment. It also measures pairwise preference accuracy to simulate human choice between generated videos. Use when the user wants to benchmark on VideoFeedback-test, GenAI-Bench, VBench, EvalCrafter, or asks about evaluating this task. Reports Spearman's ρ.
Evaluates large video language models on their ability to perceive visual details and perform multi-step reasoning over video content. It measures how well models decompose video understanding into distinct perception and reasoning stages across multiple benchmarks. Use when the user wants to benchmark on VSI-Bench, VideoMMMU, MMVU, VCR, MV, TempCom, VideoMME, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal foundation models on image and video understanding across multiple dimensions, including document and chart text recognition, mathematical reasoning, multi-image comprehension, general knowledge QA, long-form video comprehension, and temporal reasoning. Use when the user wants to benchmark on ChartQA, DocVQA, MathVista, VideoMME, Charades-STA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large video-language models for intrinsic and extrinsic hallucinations by presenting paired basic and adversarially modified Yes/No questions about video content. It probes whether models can correctly identify factual content while resisting fabricated or unverifiable details, and measures susceptibility to language bias. Use when the user wants to benchmark on VideoHallucer, or asks about evaluating this task. Reports Overall Accuracy.
Evaluates a video editing model's ability to perform multi-grained edits (class, instance, and part levels) on video-text pairs. It measures semantic alignment, temporal consistency, and pixel-level fidelity of the generated edits. Use when the user wants to benchmark on VideoGrain Evaluation Set, or asks about evaluating this task. Reports CLIP-T.
Evaluates vision-language models on video game quality assurance tasks, including glitch detection, temporal reasoning, and bug reporting. It probes the model's ability to process sampled video frames, identify visual anomalies, and generate structured or descriptive reports about game glitches. Use when the user wants to benchmark on VideoGameQA-Bench, or asks about evaluating this task. Reports accuracy.
Probes vision-language models' ability to understand video game contexts from screenshots, including recognizing actions, characters, UI elements, spatial relationships, and game mechanics. It evaluates how instruction-tuning on game-specific data improves performance compared to larger general-purpose models. Use when the user wants to benchmark on VideoGameBunny Dataset, or asks about evaluating this task. Reports performance.
Evaluates text-to-video diffusion models on visual quality and semantic alignment with input prompts. It measures intra-frame fidelity, aesthetic appeal, and inter-frame temporal consistency using automated benchmarks and human-preference predictors. Use when the user wants to benchmark on VBench, or asks about evaluating this task. Reports VBench.