Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,169–3,192 of 22,870 skills
Evaluates an AI agent's ability to maintain conversational context, resolve co-references, and ground follow-up questions in visual content. The task requires ranking a set of candidate answers based on an image and dialog history. Use when the user wants to benchmark on VisDial v0.9, or asks about evaluating this task. Reports MRR.
Evaluates the robustness of multimodal large language models (MLLMs) against vision-centric jailbreak attacks that inject realistic, image-driven contextual dialogues to elicit harmful responses. It probes safety alignment under adversarial multimodal prompts designed to bypass safety filters through semantic alignment and toxicity obfuscation. Use when the user wants to benchmark on MM-SafetyBench, SafeBench-Tiny, HarmBench, or asks about evaluating this task. Reports ASR.
Evaluates the robustness of traffic sign recognition models against adversarial attacks (PGD) and distribution shifts (ImageNet-C corruptions, color quantization). It specifically probes multi-task learning models for spurious correlations across visual attributes (color, shape, symbol, text) by measuring error propagation and task-dependent vulnerability. Use when the user wants to benchmark on VISAT, or asks about evaluating this task. Reports epsilon (model error).
Evaluates industrial anomaly detection and segmentation capabilities by measuring how well self-supervised pre-training methods transfer to identifying surface defects. It probes a model's ability to localize fine-grained anomalies in highly imbalanced, high-resolution industrial imagery under both one-class and few-shot supervised regimes. Use when the user wants to benchmark on VisA, MVTec-AD, or asks about evaluating this task. Reports AU-PR.
Evaluates computational tools and genomic features for predicting prokaryotic virus-host interactions. It probes the ability of models to correctly link viral sequences to their host taxa using either pairwise link prediction or taxonomic classification formulations. Use when the user wants to benchmark on RefSeq-VHDB, MetaHiC-VHDB, or asks about evaluating this task. Reports Top-1 accuracy.
Evaluates multimodal mathematical reasoning and high-resolution visual perception capabilities of vision-language models. It probes the model's ability to decompose complex problems into structured reasoning chunks, interleave visual tool calls, and produce accurate final answers across geometric, mathematical, and fine-grained visual benchmarks. Use when the user wants to benchmark on GeoQA, MathVista-Math, MMStar-Math, VisualProbe, V*, HR-Bench, or asks about evaluating this task. Reports a...
Evaluates a model's ability to answer clinical questions about chest X-rays and localize lesions via bounding boxes. It probes visual question answering, spatial grounding, and multi-task learning in a medical imaging context. Use when the user wants to benchmark on VinDr-CXR-VQA, or asks about evaluating this task. Reports F1 score.
Probes a model's ability to generate executable, visually faithful code (HTML, SVG, LaTeX, SMILES) from input images across diverse domains. It evaluates both syntactic correctness via execution rate and perceptual alignment with target images using coarse-to-fine visual similarity metrics. Use when the user wants to benchmark on ChartMimic, Design2Code, UniSVG, Image2Struct, Cosyn-400k, or asks about evaluating this task. Reports UniSVG Final Score.
Evaluates video language models on multilingual, culturally-diverse video understanding across 14 languages and 15 domains. It probes the models' ability to answer multiple-choice and open-ended questions about short, medium, and long videos, with a specific focus on low-resource languages and cultural reasoning. Use when the user wants to benchmark on ViMUL-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to predict the helpfulness of product reviews by jointly processing textual descriptions and visual content. It probes multimodal alignment and ranking capabilities in low-resource language settings, specifically Vietnamese. Use when the user wants to benchmark on ViMRHP, or asks about evaluating this task. Reports NDCG@K.
Evaluates Vietnamese machine reading comprehension models on span extraction from passages, including handling unanswerable questions. It measures how well systems can locate exact answer spans or correctly identify when no answer exists in the context. Use when the user wants to benchmark on UIT-ViQuAD 2.0 (ViMRC), or asks about evaluating this task. Reports F1-score.
This benchmark evaluates automatic speech recognition (ASR) models on Vietnamese medical audio containing embedded English terminology. It specifically probes the model's ability to accurately transcribe both the matrix language and code-switched segments, measuring overall transcription quality alongside specialized metrics for code-switched and non-code-switched spans. Use when the user wants to benchmark on ViMedCSS, or asks about evaluating this task. Reports WER.
This evaluation probes a model's ability to generalise to novel robotic manipulation tasks by testing robustness to instruction variations and increased task difficulty. It specifically measures compositional generalisation capabilities across four systematicity levels, ranging from object pose sensitivity to entirely novel objects and tasks. Use when the user wants to benchmark on VIMABench, or asks about evaluating this task. Reports compositional generalisation capabilities.
Evaluates video-language continual learning by testing a model's ability to retain episodic memories across streaming, long-duration videos without catastrophic forgetting. It probes cross-modal inference and temporal localization across three non-classification tasks: moment queries, natural language queries, and visual queries. Use when the user wants to benchmark on ViLCo-Bench, or asks about evaluating this task. Reports Average Recall@k (IoU=m).
Evaluates vision-language models' ability to solve multi-hop visual reasoning tasks by measuring how accurately their final predicted answers match the ground truth. It specifically probes the model's capacity for structured reasoning and answer extraction in complex domains like geometry, science, and visual question answering. Use when the user wants to benchmark on MAVIS-Geometry, A-OKVQA, GeoQA170K, CLEVR-Math, ScienceQA, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to extract robust visual features from single-channel mammography images for binary classification of malignant versus benign breast lumps. It tests cross-dataset generalization and the effectiveness of multimodal contrastive pretraining on pathological classification tasks. Use when the user wants to benchmark on MVKL, CBIS-DDSM, INbreast, or asks about evaluating this task. Reports AUC.
Evaluates a model's ability to detect hate and offensive text spans within Vietnamese social media comments. It probes sequence tagging capabilities, specifically requiring precise boundary identification of offensive content in noisy, informal text. Use when the user wants to benchmark on ViHOS, or asks about evaluating this task. Reports macro-average F1-score.
This evaluation probes a model's ability to perform scene change detection conditioned on natural language prompts. It measures how well the model distinguishes relevant semantic changes from nuisance variations across diverse domains (street-view, satellite, indoor) and handles viewpoint misalignments. Use when the user wants to benchmark on CSeg, PSCD, SYSU-CD, VL-CMU-CD, or asks about evaluating this task. Reports IoU.
Identifies and categorizes abusive content spans within long-form Vietnamese narrative texts. It probes a model's ability to perform sequence labeling for both span detection and fine-grained abuse classification across six distinct categories. Use when the user wants to benchmark on Vietnamese Narrative Abusive Span Dataset, or asks about evaluating this task. Reports Strict F-score.
Evaluates abstractive summarization capabilities on real-world and simulated medical conversations in Vietnamese, testing both human-transcribed and ASR-generated noisy transcripts. Use when the user wants to benchmark on VietMed-Sum, or asks about evaluating this task. Reports ROUGE.
Evaluates the ability of NER models to identify and classify medically defined entity spans in Vietnamese spoken text. It specifically probes robustness to ASR-generated noise and compares monolingual vs. multilingual, encoder vs. seq2seq architectures. Use when the user wants to benchmark on VietMed-NER, or asks about evaluating this task. Reports micro F1 score.
Evaluates automatic speech recognition (ASR) performance on Vietnamese medical domain audio. It measures how well models transcribe speech containing medical terminology and regional accents, assessing cross-domain transfer capabilities. Use when the user wants to benchmark on VietMed, or asks about evaluating this task. Reports WER.
This benchmark evaluates a model's ability to generate culturally accurate and coherent explanations for Vietnamese visual question answering. It probes both linguistic fluency and the model's capacity to ground visual evidence in domain-specific cultural knowledge through structured, stepwise reasoning. Use when the user wants to benchmark on Vietnamese VQA dataset, or asks about evaluating this task. Reports Cultural Accuracy.
This benchmark evaluates end-to-end Retrieval Augmented Generation (RAG) systems on visually rich, real-world documents across multiple professional domains. It probes a model's ability to retrieve relevant pages, generate accurate answers to complex open-ended and multi-hop queries, and precisely ground those answers with bounding boxes in multimodal content. Use when the user wants to benchmark on ViDoRe V3, or asks about evaluating this task. Reports F1 score (Dice coefficient).