Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,870
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,1453,168 of 22,870 skills

Visual Spatial Reasoning EvalA

This benchmark evaluates visual language models' ability to understand and reason about spatial relationships between objects in images. It specifically probes orientation-dependent relations, frame-of-reference shifts (intrinsic vs. relative), and zero-shot generalization to unseen object concepts. Use when the user wants to benchmark on VSR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Semantic Segmentation EvalA

Evaluates a model's ability to assign a semantic class label to every pixel in an image, capturing fine-grained scene understanding. It measures pixel-level classification accuracy and boundary alignment across diverse outdoor and indoor environments. Use when the user wants to benchmark on Pascal Context, Sift Flow, COCO Stuff, or asks about evaluating this task. Reports GPA.

researchpythonperformance
0
3
Visual Rl Rectification EvalA

Evaluates the stability and generalization of PPO policies trained with mode-dependent layers (BatchNorm, dropout) across visual reinforcement learning environments. It probes whether a deterministic rectification phase prevents reward collapse and aligns training-evaluation dynamics compared to standard training modes. Use when the user wants to benchmark on Procgen, Histopathology Patch-Localization, Natural Image Patch-Localization, or asks about evaluating this task. Reports normalized re...

researchpythonexpress
0
3
Visual Relationship Detection EvalA

Probes a model's ability to identify and classify interactions between pairs of objects in an image (subject-predicate-object triples). It focuses on capturing relational semantics beyond isolated object detection. Use when the user wants to benchmark on Visual Relationship Dataset, Visual Genome, or asks about evaluating this task. Reports recall@50.

researchpythongo
0
3
Visual Reasoning EvalA

Evaluates vision-language models on complex visual reasoning tasks including arithmetic counting, structural perception, and spatial transformations. It specifically probes the model's ability to generalize under domain shifts and adapt to distribution changes with limited data. Use when the user wants to benchmark on CLEVR-Math, Super-CLEVR, Geo170K/Math360K/Geometry3K, TRANCE, or asks about evaluating this task. Reports accuracy-rate (Acc).

researchpythongo
0
3
Visual Quality Assessment EvalA

Evaluates the perceptual visual quality of interpolated frames generated by optical flow methods against human judgments. It measures how well traditional objective metrics like RMSE correlate with crowdsourced subjective quality ratings across multiple video sequences. Use when the user wants to benchmark on Middlebury, or asks about evaluating this task. Reports SROCC.

researchpythonrust
0
3
Visual Prompt EvalA

Evaluates multimodal large language models' ability to comprehend and reason about visual prompts (points, bounding boxes, free-form shapes) for fine-grained object classification, region captioning, OCR, and complex visual reasoning. Use when the user wants to benchmark on LVIS, PACO, COCO-Text, RefCOCOg, MDVP-Bench, LLaVA-Bench, Ferret-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Visual Information Extraction EvalA

Evaluates a model's ability to extract entity spans and link them to key-value pairs from complex, real-world document images. It probes joint vision-language understanding, handling poor image quality, occlusion, and multi-lingual text without relying on external OCR pipelines. Use when the user wants to benchmark on FUNSD, XFUND, CORD, SIBR, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Visual Genome Sgg EvalA

Evaluates fine-grained scene graph generation by predicting subject-predicate-object triplets from images. It measures recall and F1 scores across head, body, and tail predicate classes to assess performance on long-tailed distributions and missing annotations. Use when the user wants to benchmark on Visual Genome, or asks about evaluating this task. Reports mR@K, F@K.

researchpythongo
0
3
Visual Counterfact EvalA

Evaluates how vision-language models resolve conflicts between visual input and language priors by reasoning about altered visual attributes (color and size). It probes whether models rely on visual evidence or textual priors when they contradict. Use when the user wants to benchmark on Visual-Counterfact, or asks about evaluating this task. Reports MAC.

researchpythongo
0
3
Visual Cot EvalA

Evaluates multi-modal large language models' ability to perform chain-of-thought reasoning with dynamic visual focusing on specific image regions. It probes localized visual understanding, intermediate bounding box prediction, and multi-turn reasoning across document, chart, general VQA, relation reasoning, and fine-grained domains. Use when the user wants to benchmark on Visual CoT Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visual Commonsense EvalA

Evaluates language models' zero-shot visual and textual commonsense reasoning without relying on ground-truth images. It probes the model's ability to infer object properties (color, shape, size) and answer general knowledge questions by internally generating and fusing multiple image variations from text prompts. Use when the user wants to benchmark on ImageNetVC, Object Commonsense (Memory Color, Color Terms, ViComTe, Size), Commonsense Reasoning (PIQA, SIQA, HellaSwag, WinoGrande, ARC, Ope...

researchpythongo
0
3
Vista Score EvalA

Evaluates conversational factuality and hallucination detection in LLMs by decomposing dialogue turns into atomic claims, verifying them against reference texts and dialogue history, and categorizing unverifiable content. It measures how well models track factual consistency across sequential turns. Use when the user wants to benchmark on FaithDial, or asks about evaluating this task. Reports claim-level accuracy.

researchpythongo
0
3
Vista Multimodal EvalA

Evaluates cross-modal vision-text alignment in Multimodal Large Language Models (MLLMs) across high-level semantic VQA, general multimodal understanding, and fine-grained visual perception/retrieval tasks. Use when the user wants to benchmark on VQAv2, OK-VQA, GQA, TextVQA, RealWorldQA, DocVQA, MMBench, SEED, AI2D, MMMU, MMStar, MME, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports performance.

researchpythongo
0
3
Vista EvalA

This benchmark evaluates a model's ability to generate concise, structured summaries of scientific conference talks from video inputs. It specifically probes informativeness, alignment with visual/audio content, and factual consistency against the corresponding paper abstracts. Use when the user wants to benchmark on VISTA, or asks about evaluating this task. Reports ROUGE-1 F1.

researchpython
0
3
Visres Bench EvalA

This benchmark evaluates the visual reasoning capabilities of vision-language models across a perceptual-to-reasoning continuum. It isolates three levels of difficulty: basic perceptual grounding under transformations, single-attribute reasoning (color, count, orientation), and multi-attribute compositional reasoning. The setup tests whether models rely on genuine visual abstraction or fall back to linguistic priors when faced with naturalistic perturbations and rule-based inference. Use when...

researchpythongo
0
3
Visquic Http3 Response Estimation EvalA

Evaluates a model's ability to estimate the number of HTTP/3 responses in encrypted QUIC traffic using only observable packet characteristics. It probes the capability to extract meaningful temporal and structural patterns from encrypted flows without plaintext inspection. Use when the user wants to benchmark on VisQUIC, or asks about evaluating this task. Reports CAP±k.

researchpythongit
0
3
VisorA

Evaluates whether text-to-image models correctly render spatial relationships between objects mentioned in a prompt. It disentangles object generation accuracy from spatial correctness to reveal model biases like object priority and merging. Use when the user has predictions and gold and needs to compute VISOR.

researchpythongo
0
3
Visonlyqa EvalA

This benchmark probes a model's ability to accurately perceive basic geometric information—such as shape, angle, length, area, and intersections—in scientific figures and diagrams. It isolates visual perception from higher-level reasoning or domain knowledge by using direct, low-reasoning questions on synthetic and real-world images. Use when the user wants to benchmark on VisOnlyQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Visnumbench EvalA

Evaluates the intuitive number sense of Multimodal Large Language Models (MLLMs) by testing their ability to estimate and reason about seven visual numerical attributes (angle, scale, length, quantity, depth, area, volume) across four estimation tasks (range estimation, value estimation, value comparison, multiplicative estimation). Use when the user wants to benchmark on VisNumBench, or asks about evaluating this task. Reports Accuracy (%).

researchpythongo
0
3
Visit Bench EvalA

Evaluates vision-language models' ability to follow complex, real-world instructions on images. It probes open-ended generation, context-sensitive reasoning, and instruction-conditioned captioning by measuring how well model outputs align with human preferences and high-quality references. Use when the user wants to benchmark on VisIT-Bench, or asks about evaluating this task. Reports Elo rating, Win rate vs. reference.

researchpythongo
0
3
Vision R1 EvalA

Evaluates Large Vision-Language Models on their ability to detect, localize, and ground objects in images across diverse and challenging scenarios, including in-domain dense detection, out-of-domain real-world settings, and generalization to unseen categories or scenes. Use when the user wants to benchmark on MSCOCO Val2017, ODINW-13, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Vision Language Ood EvalA

Probes the ability of vision-language models to distinguish in-distribution from out-of-distribution samples under semantic, covariate, and real-world distribution shifts. It evaluates both zero-shot and few-shot prompt learning approaches across multiple benchmarks to assess robustness and ranking consistency. Use when the user wants to benchmark on ImageNet-X, ImageNet-FS-X, Wilds-FS-X, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Vision Arch Gen EvalA

Evaluates the classification performance of LLM-generated neural network architectures by training each for a single epoch on seven computer vision benchmarks. It measures Top-1 accuracy to assess architectural quality, while also tracking generation efficiency via hash validation speed and duplicate rejection rates. Use when the user wants to benchmark on MNIST, CelebA-Gender, CIFAR-10, CIFAR-100, ImageNette, SVHN, Places365, or asks about evaluating this task. Reports Top-1 accuracy after 1...

researchpythonapi
0
3