Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 3,121–3,144 of 22,870 skills
Evaluates the robustness of Vision-Language Models against Gaussian noise perturbations on input images, measuring both capability degradation (helpfulness, OCR, knowledge) and safety alignment (toxicity, attack success rate) under noisy conditions. It probes whether noise-augmented fine-tuning preserves model utility while mitigating vulnerability to adversarial or distribution-shifted visual inputs. Use when the user wants to benchmark on MM-Vet, RealToxicityPrompts, or asks about evaluatin...
Evaluates vision-language models on instruction-following and multimodal reasoning tasks across multiple established benchmarks. Probes capabilities in general VQA, mathematical reasoning, scientific understanding, hallucination detection, and multilingual comprehension. Use when the user wants to benchmark on MMBench, MME, MathVista, HallusionBench, SEEDBench, LLaVABench, ScienceQA, or asks about evaluating this task. Reports evaluation metric.
Evaluates Vietnamese large language models on contextual reasoning, academic knowledge, general trivia, and long-form reading comprehension. Probes both language modeling capability (perplexity) and factual/reasoning accuracy across culturally and linguistically specific tasks. Use when the user wants to benchmark on LAMBADA Vietnamese, Exam Vietnamese, General Knowledge, Comprehension QA, or asks about evaluating this task. Reports accuracy.
Evaluates the safety alignment and helpfulness of vision-language models (VLLMs) by measuring their ability to reject harmful image-text prompts while maintaining performance on benign queries. Use when the user wants to benchmark on VLGuard, or asks about evaluating this task. Reports ASR.
Evaluates large language models on Vietnamese legal reasoning within a civil law framework. It probes capabilities ranging from statutory recall and hierarchical navigation to multi-step conflict detection, penalty estimation, and ethical bias analysis. Use when the user wants to benchmark on VLegal-Bench, or asks about evaluating this task. Reports Accuracy.
This evaluation probes a model's embodied reasoning capabilities, including spatial understanding, visual grounding, task planning, and closed-loop robotic control. It measures how well vision-language models transfer general multimodal knowledge to robot-specific manipulation tasks and identifies the domain gap between internet-scale pretraining and real-world embodiment. Use when the user wants to benchmark on ERQA, Ego-Plan2, Where2place, Pointarena, Paco-Lavis, Pixmo-Points, VSI-Bench, Re...
Evaluates the generalization, long-horizon reasoning, and language-conditioned manipulation capabilities of Vision-Language-Action (VLA) models, workflow frameworks, and Vision-Language Models (VLMs) in simulated robotic environments. It probes performance across seen/unseen objects, semantic instruction understanding, and composite task decomposition. Use when the user wants to benchmark on VLABench, or asks about evaluating this task. Reports task_progress_score.
Evaluates a vision-language-action model's ability to generalize across diverse robotic embodiments, simulation environments, and real-world platforms. It probes cross-embodiment adaptation, parameter-efficient fine-tuning capabilities, and dexterous manipulation performance. Use when the user wants to benchmark on Libero, Simpler, Calvin, VLABench, RoboTwin-2.0, NAVSIM, BridgeData-v2, Soft-Fold, or asks about evaluating this task. Reports success_rate.
Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.
Evaluates the multimodal reasoning and self-reflection capabilities of vision-language models across math, multi-discipline, and real-world benchmarks. It probes whether models can correctly interpret visual-textual inputs and produce accurate final answers under greedy decoding. Use when the user wants to benchmark on MathVista, MathVerse, MathVision, MMMU-Pro, MMMU, EMMA, MegaBench, or asks about evaluating this task. Reports Pass@1 accuracy.
Evaluates zero-shot video understanding capabilities on action recognition and text-to-video retrieval tasks across diverse benchmarks. Use when the user wants to benchmark on Something-something-v2 (SSv2), EPIC-KITCHENS-100 (EK-100), EgoExo4D Keysteps, Kinetics-400, COIN, CrossTask, MSR-VTT, ActivityNet, DiDeMo, MSVD, YouCook2, PVD-Bench, Dream-1K, VDC-1K, or asks about evaluating this task. Reports top-1 accuracy, recall@1.
Evaluates vision-language models on compositional reasoning capabilities, specifically testing their ability to correctly bind attributes, understand semantic relations, and parse word order in image-text pairs. It also measures systematic generalization to unseen concept combinations and zero-shot classification and retrieval performance. Use when the user wants to benchmark on ARO, CREPE, SVO, VL-Checklist, or asks about evaluating this task. Reports accuracy.
Evaluates video editing models on local, entity-level modifications (addition, modification, deletion) by measuring background preservation, text alignment, temporal consistency, and visual quality. Use when the user wants to benchmark on VIVID-10M-Eval, or asks about evaluating this task. Reports Text Alignment (TA).
Evaluates long-term multivariate time-series forecasting of intraoperative vital signs under three clinically realistic conditions: complete data, variable missingness, and cross-center generalization. It probes a model's ability to handle heterogeneous clinical data, adapt to missing sensor inputs, and generalize across different hospital centers. Use when the user wants to benchmark on VitalDB, MOVER-SIS, or asks about evaluating this task. Reports MAE.
Evaluates the ability of Vision Transformer models combined with dimensionality reduction and clustering algorithms to perform zero-shot species-level clustering of animal images. It probes how well unsupervised pipelines can recover ground-truth taxonomic labels and capture intra-specific variation without manual annotation. Use when the user wants to benchmark on Animal Images (Birds & Mammals), or asks about evaluating this task. Reports V-measure.
Evaluates the robustness of Vision Transformer models against input perturbations including adversarial attacks (FGSM/PGD), spatial transformations, and restricted attention. It probes whether ViTs maintain classification performance under distribution shifts and targeted attacks compared to standard CNNs. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models' ability to perform abstract visual reasoning across five fine-grained perceptual dimensions (numerosity, attributes, style, position, spatial relations) and two high-level reasoning tasks (analogical pattern matching and constraint-based logic). Use when the user wants to benchmark on VisuRiddles, or asks about evaluating this task. Reports exact match.
VisuLogic probes vision-centric reasoning in multimodal large language models by presenting problems that require retaining critical visual cues during image description. It eliminates text-based reasoning shortcuts, forcing models to perform genuine visual inference across categories like spatial relations, quantitative shifts, and stylistic details. Use when the user wants to benchmark on VisuLogic, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.
This benchmark probes fine-grained visual understanding of Vision-Language Models in densely populated, high-resolution scenes. It evaluates capabilities across six core tasks including activity recognition, attribute recognition, counting, OCR, visual reasoning, and global scene classification. Use when the user wants to benchmark on VisualOverload, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates deep learning models on fine-grained bird species classification and spatio-temporal behavior recognition in ecological video footage. It probes the model's ability to localize birds, identify their species, and classify their actions across video frames in real-world wetland environments. Use when the user wants to benchmark on Visual WetlandBirds Dataset, or asks about evaluating this task. Reports accuracy.
Evaluates multimodal large language models' ability to perform precise spatial reasoning and visual text grounding in document images. It tests whether models can generate accurate bounding boxes that support their textual answers, both from scratch (OCR-free) and when provided with OCR text (OCR-based), while also measuring their instruction-following capability. Use when the user wants to benchmark on ChartQA, DocVQA, InfographicsVQA, TRINS, or asks about evaluating this task. Reports IoU.
Probes multimodal visual reasoning capabilities over complex, LaTeX-rendered table images. It specifically tests multi-step inference, structural layout understanding, and the ability to extract and reason over tabular data from visual inputs rather than raw text. Use when the user wants to benchmark on Visual-TableQA, or asks about evaluating this task. Reports Relaxed Accuracy.
This evaluation probes how Vision-Language Models ground their responses in visual input versus relying on language priors or user bias. It measures perceptual awareness, visual dependency, and alignment conflicts by comparing model behavior across original, blank, noisy, and semantically conflicting images. Use when the user wants to benchmark on GQA, VQAv2, A-OKVQA, POPE, or asks about evaluating this task. Reports VNS.