Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,479
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,745–3,768 of 23,479 skills

Vit Robustness EvalA

Evaluates the robustness of Vision Transformer models against input perturbations including adversarial attacks (FGSM/PGD), spatial transformations, and restricted attention. It probes whether ViTs maintain classification performance under distribution shifts and targeted attacks compared to standard CNNs. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Visuriddles EvalA

Evaluates multimodal large language models' ability to perform abstract visual reasoning across five fine-grained perceptual dimensions (numerosity, attributes, style, position, spatial relations) and two high-level reasoning tasks (analogical pattern matching and constraint-based logic). Use when the user wants to benchmark on VisuRiddles, or asks about evaluating this task. Reports exact match.

researchpythongo
0
3
Visulogic EvalA

VisuLogic probes vision-centric reasoning in multimodal large language models by presenting problems that require retaining critical visual cues during image description. It eliminates text-based reasoning shortcuts, forcing models to perform genuine visual inference across categories like spatial relations, quantitative shifts, and stylistic details. Use when the user wants to benchmark on VisuLogic, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visualwebbench EvalA

Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.

researchpythongo
0
3
Visualoverload EvalA

This benchmark probes fine-grained visual understanding of Vision-Language Models in densely populated, high-resolution scenes. It evaluates capabilities across six core tasks including activity recognition, attribute recognition, counting, OCR, visual reasoning, and global scene classification. Use when the user wants to benchmark on VisualOverload, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Wetlandbirds EvalA

This benchmark evaluates deep learning models on fine-grained bird species classification and spatio-temporal behavior recognition in ecological video footage. It probes the model's ability to localize birds, identify their species, and classify their actions across video frames in real-world wetland environments. Use when the user wants to benchmark on Visual WetlandBirds Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Text Grounding EvalA

Evaluates multimodal large language models' ability to perform precise spatial reasoning and visual text grounding in document images. It tests whether models can generate accurate bounding boxes that support their textual answers, both from scratch (OCR-free) and when provided with OCR text (OCR-based), while also measuring their instruction-following capability. Use when the user wants to benchmark on ChartQA, DocVQA, InfographicsVQA, TRINS, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Visual Tableqa EvalA

Probes multimodal visual reasoning capabilities over complex, LaTeX-rendered table images. It specifically tests multi-step inference, structural layout understanding, and the ability to extract and reason over tabular data from visual inputs rather than raw text. Use when the user wants to benchmark on Visual-TableQA, or asks about evaluating this task. Reports Relaxed Accuracy.

researchpythongo
0
3
Visual Sycophancy EvalA

This evaluation probes how Vision-Language Models ground their responses in visual input versus relying on language priors or user bias. It measures perceptual awareness, visual dependency, and alignment conflicts by comparing model behavior across original, blank, noisy, and semantically conflicting images. Use when the user wants to benchmark on GQA, VQAv2, A-OKVQA, POPE, or asks about evaluating this task. Reports VNS.

researchpythongo
0
3
Visual Spatial Reasoning EvalA

This benchmark evaluates visual language models' ability to understand and reason about spatial relationships between objects in images. It specifically probes orientation-dependent relations, frame-of-reference shifts (intrinsic vs. relative), and zero-shot generalization to unseen object concepts. Use when the user wants to benchmark on VSR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Semantic Segmentation EvalA

Evaluates a model's ability to assign a semantic class label to every pixel in an image, capturing fine-grained scene understanding. It measures pixel-level classification accuracy and boundary alignment across diverse outdoor and indoor environments. Use when the user wants to benchmark on Pascal Context, Sift Flow, COCO Stuff, or asks about evaluating this task. Reports GPA.

researchpythonperformance
0
3
Visual Rl Rectification EvalA

Evaluates the stability and generalization of PPO policies trained with mode-dependent layers (BatchNorm, dropout) across visual reinforcement learning environments. It probes whether a deterministic rectification phase prevents reward collapse and aligns training-evaluation dynamics compared to standard training modes. Use when the user wants to benchmark on Procgen, Histopathology Patch-Localization, Natural Image Patch-Localization, or asks about evaluating this task. Reports normalized re...

researchpythonexpress
0
3
Visual Relationship Detection EvalA

Probes a model's ability to identify and classify interactions between pairs of objects in an image (subject-predicate-object triples). It focuses on capturing relational semantics beyond isolated object detection. Use when the user wants to benchmark on Visual Relationship Dataset, Visual Genome, or asks about evaluating this task. Reports recall@50.

researchpythongo
0
3
Visual Reasoning EvalA

Evaluates vision-language models on complex visual reasoning tasks including arithmetic counting, structural perception, and spatial transformations. It specifically probes the model's ability to generalize under domain shifts and adapt to distribution changes with limited data. Use when the user wants to benchmark on CLEVR-Math, Super-CLEVR, Geo170K/Math360K/Geometry3K, TRANCE, or asks about evaluating this task. Reports accuracy-rate (Acc).

researchpythongo
0
3
Visual Quality Assessment EvalA

Evaluates the perceptual visual quality of interpolated frames generated by optical flow methods against human judgments. It measures how well traditional objective metrics like RMSE correlate with crowdsourced subjective quality ratings across multiple video sequences. Use when the user wants to benchmark on Middlebury, or asks about evaluating this task. Reports SROCC.

researchpythonrust
0
3
Visual Prompt EvalA

Evaluates multimodal large language models' ability to comprehend and reason about visual prompts (points, bounding boxes, free-form shapes) for fine-grained object classification, region captioning, OCR, and complex visual reasoning. Use when the user wants to benchmark on LVIS, PACO, COCO-Text, RefCOCOg, MDVP-Bench, LLaVA-Bench, Ferret-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Visual Information Extraction EvalA

Evaluates a model's ability to extract entity spans and link them to key-value pairs from complex, real-world document images. It probes joint vision-language understanding, handling poor image quality, occlusion, and multi-lingual text without relying on external OCR pipelines. Use when the user wants to benchmark on FUNSD, XFUND, CORD, SIBR, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Visual Genome Sgg EvalA

Evaluates fine-grained scene graph generation by predicting subject-predicate-object triplets from images. It measures recall and F1 scores across head, body, and tail predicate classes to assess performance on long-tailed distributions and missing annotations. Use when the user wants to benchmark on Visual Genome, or asks about evaluating this task. Reports mR@K, F@K.

researchpythongo
0
3
Visual Counterfact EvalA

Evaluates how vision-language models resolve conflicts between visual input and language priors by reasoning about altered visual attributes (color and size). It probes whether models rely on visual evidence or textual priors when they contradict. Use when the user wants to benchmark on Visual-Counterfact, or asks about evaluating this task. Reports MAC.

researchpythongo
0
3
Visual Cot EvalA

Evaluates multi-modal large language models' ability to perform chain-of-thought reasoning with dynamic visual focusing on specific image regions. It probes localized visual understanding, intermediate bounding box prediction, and multi-turn reasoning across document, chart, general VQA, relation reasoning, and fine-grained domains. Use when the user wants to benchmark on Visual CoT Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Commonsense EvalA

Evaluates language models' zero-shot visual and textual commonsense reasoning without relying on ground-truth images. It probes the model's ability to infer object properties (color, shape, size) and answer general knowledge questions by internally generating and fusing multiple image variations from text prompts. Use when the user wants to benchmark on ImageNetVC, Object Commonsense (Memory Color, Color Terms, ViComTe, Size), Commonsense Reasoning (PIQA, SIQA, HellaSwag, WinoGrande, ARC, Ope...

researchpythongo
0
3
Vista Score EvalA

Evaluates conversational factuality and hallucination detection in LLMs by decomposing dialogue turns into atomic claims, verifying them against reference texts and dialogue history, and categorizing unverifiable content. It measures how well models track factual consistency across sequential turns. Use when the user wants to benchmark on FaithDial, or asks about evaluating this task. Reports claim-level accuracy.

researchpythongo
0
3
Vista Multimodal EvalA

Evaluates cross-modal vision-text alignment in Multimodal Large Language Models (MLLMs) across high-level semantic VQA, general multimodal understanding, and fine-grained visual perception/retrieval tasks. Use when the user wants to benchmark on VQAv2, OK-VQA, GQA, TextVQA, RealWorldQA, DocVQA, MMBench, SEED, AI2D, MMMU, MMStar, MME, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Vista EvalA

This benchmark evaluates a model's ability to generate concise, structured summaries of scientific conference talks from video inputs. It specifically probes informativeness, alignment with visual/audio content, and factual consistency against the corresponding paper abstracts. Use when the user wants to benchmark on VISTA, or asks about evaluating this task. Reports ROUGE-1 F1.

researchpython
0
3