Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,479
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,721–3,744 of 23,479 skills

Vn Mteb EvalA

Evaluates the quality of Vietnamese text embeddings across six standard information retrieval and NLP tasks. It probes a model's ability to capture semantic similarity, perform document retrieval, classify text, cluster documents, and rank relevant passages in Vietnamese. Use when the user wants to benchmark on VN-MTEB, or asks about evaluating this task. Reports Average Task Score.

researchpythongo
0
3
Vmbench EvalA

Evaluates how well text-to-video models generate videos that align with human perceptual preferences across five motion dimensions: object integrity, motion smoothness, commonsense adherence, perceptible amplitude, and temporal coherence. Use when the user wants to benchmark on MMPG-set, or asks about evaluating this task. Reports Spearman correlation.

researchpythongo
0
3
Vlue EvalA

Evaluates Vietnamese natural language understanding across five diverse tasks including machine reading comprehension, natural language inference, emotion recognition, hate speech detection, and part-of-speech tagging. It assesses a model's ability to comprehend text, reason over sentence pairs, classify emotions and hate speech, and perform syntactic analysis in Vietnamese. Use when the user wants to benchmark on UIT-ViQuAD 2.0, ViNLI, VSMEC, ViHOS, NIIVTB POS, or asks about evaluating this ...

researchpythongo
0
3
Vlsbench EvalA

Evaluates the safety alignment of multimodal large language models (MLLMs) by testing their ability to correctly identify and appropriately respond to unsafe image-text pairs. It specifically probes how well models handle Visual Safety Information Leakage (VSIL), where harmful content might be implicitly revealed in the textual query rather than the image. Use when the user wants to benchmark on VLSBench, or asks about evaluating this task. Reports safety rate (%).

researchpythongo
0
3
Vln Task Planning EvalA

Evaluates an agent's ability to decompose coarse-grained natural language navigation instructions into executable subtasks and navigate through simulated environments to reach target locations or interact with objects. It probes task planning, visual-language grounding, and dynamic error recovery in continuous or discrete navigation spaces. Use when the user wants to benchmark on R2R, REVERIE, ALFRED, or asks about evaluating this task. Reports Success Rate (SR).

researchpythongo
0
3
Vln Ce EvalA

Evaluates an agent's ability to follow natural language instructions to navigate to a target location in a continuous 3D environment. It probes low-level action control, obstacle avoidance, and spatial reasoning without relying on a pre-defined graph topology or oracle localization. Use when the user wants to benchmark on VLN-CE, or asks about evaluating this task. Reports SR, SPL.

researchpythongo
0
3
Vlmbench EvalA

This benchmark evaluates a robot agent's ability to execute 6D manipulation tasks guided by natural language instructions and visual observations. It probes compositional reasoning, object localization, and precise pose estimation in both seen and unseen object settings. Use when the user wants to benchmark on VLMbench, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Vlm Safety EvalA

Evaluates the safety and generalization capabilities of vision-language models by measuring their ability to redirect unsafe content to safe alternatives, maintain zero-shot classification accuracy, and generate safe text/images from unsafe prompts or inputs. Use when the user wants to benchmark on ViSU, NSFWCaps, I2P, NudeNet/SMID/NSFW URLs, Zero-shot Benchmarks (ImageNet variants, Caltech101, Oxford Pets, Flowers102, Stanford Cars, UCF101, DTD), or asks about evaluating this task. Reports %...

researchpythongo
0
3
Vlm Interaction Reasoning EvalA

Evaluates vision-language models on general visual understanding, spatial/relational reasoning, and specifically interactional reasoning in dynamic scenes using a suite of standard VQA and scene understanding benchmarks. Use when the user wants to benchmark on VQAv2, VizWiz, TextVQA, GQA, VSR, RealWorldQA, MMT-Bench, SEEDBench, A-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vlm Gaussian Noise Robustness EvalA

Evaluates the robustness of Vision-Language Models against Gaussian noise perturbations on input images, measuring both capability degradation (helpfulness, OCR, knowledge) and safety alignment (toxicity, attack success rate) under noisy conditions. It probes whether noise-augmented fine-tuning preserves model utility while mitigating vulnerability to adversarial or distribution-shifted visual inputs. Use when the user wants to benchmark on MM-Vet, RealToxicityPrompts, or asks about evaluatin...

researchpythongo
0
3
Vlm Benchmarks EvalA

Evaluates vision-language models on instruction-following and multimodal reasoning tasks across multiple established benchmarks. Probes capabilities in general VQA, mathematical reasoning, scientific understanding, hallucination detection, and multilingual comprehension. Use when the user wants to benchmark on MMBench, MME, MathVista, HallusionBench, SEEDBench, LLaVABench, ScienceQA, or asks about evaluating this task. Reports evaluation metric.

researchpythongo
0
3
Vllm EvalA

Evaluates Vietnamese large language models on contextual reasoning, academic knowledge, general trivia, and long-form reading comprehension. Probes both language modeling capability (perplexity) and factual/reasoning accuracy across culturally and linguistically specific tasks. Use when the user wants to benchmark on LAMBADA Vietnamese, Exam Vietnamese, General Knowledge, Comprehension QA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vlguard EvalA

Evaluates the safety alignment and helpfulness of vision-language models (VLLMs) by measuring their ability to reject harmful image-text prompts while maintaining performance on benign queries. Use when the user wants to benchmark on VLGuard, or asks about evaluating this task. Reports ASR.

researchpythongit
0
3
Vlegal Bench EvalA

Evaluates large language models on Vietnamese legal reasoning within a civil law framework. It probes capabilities ranging from statutory recall and hierarchical navigation to multi-step conflict detection, penalty estimation, and ethical bias analysis. Use when the user wants to benchmark on VLegal-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Vlaser Embodied Reasoning EvalA

This evaluation probes a model's embodied reasoning capabilities, including spatial understanding, visual grounding, task planning, and closed-loop robotic control. It measures how well vision-language models transfer general multimodal knowledge to robot-specific manipulation tasks and identifies the domain gap between internet-scale pretraining and real-world embodiment. Use when the user wants to benchmark on ERQA, Ego-Plan2, Where2place, Pointarena, Paco-Lavis, Pixmo-Points, VSI-Bench, Re...

researchpythongo
0
3
Vlabench EvalA

Evaluates the generalization, long-horizon reasoning, and language-conditioned manipulation capabilities of Vision-Language-Action (VLA) models, workflow frameworks, and Vision-Language Models (VLMs) in simulated robotic environments. It probes performance across seen/unseen objects, semantic instruction understanding, and composite task decomposition. Use when the user wants to benchmark on VLABench, or asks about evaluating this task. Reports task_progress_score.

researchpythongo
0
3
Vla Cross Embodiment EvalA

Evaluates a vision-language-action model's ability to generalize across diverse robotic embodiments, simulation environments, and real-world platforms. It probes cross-embodiment adaptation, parameter-efficient fine-tuning capabilities, and dexterous manipulation performance. Use when the user wants to benchmark on Libero, Simpler, Calvin, VLABench, RoboTwin-2.0, NAVSIM, BridgeData-v2, Soft-Fold, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Vl Rewardbench EvalA

Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.

researchpythongo
0
3
Vl Rethinker EvalA

Evaluates the multimodal reasoning and self-reflection capabilities of vision-language models across math, multi-discipline, and real-world benchmarks. It probes whether models can correctly interpret visual-textual inputs and produce accurate final answers under greedy decoding. Use when the user wants to benchmark on MathVista, MathVerse, MathVision, MMMU-Pro, MMMU, EMMA, MegaBench, or asks about evaluating this task. Reports Pass@1 accuracy.

researchpythongo
0
3
Vl Jepa Zero Shot BenchmarksA

Evaluates zero-shot video understanding capabilities on action recognition and text-to-video retrieval tasks across diverse benchmarks. Use when the user wants to benchmark on Something-something-v2 (SSv2), EPIC-KITCHENS-100 (EK-100), EgoExo4D Keysteps, Kinetics-400, COIN, CrossTask, MSR-VTT, ActivityNet, DiDeMo, MSVD, YouCook2, PVD-Bench, Dream-1K, VDC-1K, or asks about evaluating this task. Reports top-1 accuracy, recall@1.

researchpythongo
0
3
Vl Compositionality EvalA

Evaluates vision-language models on compositional reasoning capabilities, specifically testing their ability to correctly bind attributes, understand semantic relations, and parse word order in image-text pairs. It also measures systematic generalization to unseen concept combinations and zero-shot classification and retrieval performance. Use when the user wants to benchmark on ARO, CREPE, SVO, VL-Checklist, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vivd 10m EvalA

Evaluates video editing models on local, entity-level modifications (addition, modification, deletion) by measuring background preservation, text alignment, temporal consistency, and visual quality. Use when the user wants to benchmark on VIVID-10M-Eval, or asks about evaluating this task. Reports Text Alignment (TA).

researchpythonaws
0
3
Vitalbench EvalA

Evaluates long-term multivariate time-series forecasting of intraoperative vital signs under three clinically realistic conditions: complete data, variable missingness, and cross-center generalization. It probes a model's ability to handle heterogeneous clinical data, adapt to missing sensor inputs, and generalize across different hospital centers. Use when the user wants to benchmark on VitalDB, MOVER-SIS, or asks about evaluating this task. Reports MAE.

researchpythongo
0
3
Vit Zero Shot Clustering EvalA

Evaluates the ability of Vision Transformer models combined with dimensionality reduction and clustering algorithms to perform zero-shot species-level clustering of animal images. It probes how well unsupervised pipelines can recover ground-truth taxonomic labels and capture intra-specific variation without manual annotation. Use when the user wants to benchmark on Animal Images (Birds & Mammals), or asks about evaluating this task. Reports V-measure.

researchpythongo
0
3