Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,484
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,841–3,864 of 23,484 skills

Vid Ad EvalA

Probes image-level logical anomaly detection under vision-induced distractions such as background changes, blur, and low light. It tests whether models can identify violations of logical constraints (e.g., quantity, length, type, placement) by reasoning over textual descriptions rather than relying on brittle low-level visual features. Use when the user wants to benchmark on VID-AD, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Vicuna Benchmark EvalA

This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.

researchpythonsecurity
0
3
Vibepass EvalA

Evaluates LLMs on fault-targeted test generation and fault-targeted program repair. It probes discriminative fault detection, fault hypothesis generation, and the ability to debug subtle semantic bugs under diagnostic guidance. Use when the user wants to benchmark on VIBEPASS, or asks about evaluating this task. Reports D_{IO}.

researchpythondebugging
0
3
Vibe Eval EvalA

Probes multimodal reasoning and visual understanding on real-world images. Specifically designed with a 'hard' subset of prompts that are unsolvable by current frontier models to measure genuine performance gaps and contamination-free generalization. Use when the user wants to benchmark on Vibe-Eval, or asks about evaluating this task. Reports Vibe-Eval Score.

researchpythongit
0
3
Vib Probe Hallucination EvalA

Evaluates the ability of Vision-Language Models to avoid generating unfaithful or non-existent visual details in both closed-set discriminative QA and open-ended generative captioning. It also measures how well an auxiliary probing framework can detect these hallucinations via attention dynamics and mitigate them at inference time without retraining. Use when the user wants to benchmark on POPE, AMBER, M-HalDetect, COCO-Caption, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
Vhelm EvalA

Holistic evaluation of vision-language models across multiple dimensions including visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. Use when the user wants to benchmark on VHELM Scenarios, or asks about evaluating this task. Reports scenario_score.

researchpythongit
0
3
Vhd11k EvalA

Evaluates multimodal models' ability to detect harmful content in images and videos across ten specific harmful categories and a general unharmful class. It probes binary classification robustness against dataset imbalance and multi-class reasoning capabilities under varying prompt conditions. Use when the user wants to benchmark on VHD11K, SMID, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vgphrasecut EvalA

This benchmark evaluates language-based image segmentation by requiring models to ground natural language phrases into precise image regions. It probes a model's ability to handle long-tail categories, attributes, relationships, and varying object sizes in open-vocabulary settings. Use when the user wants to benchmark on VGPhraseCut, or asks about evaluating this task. Reports mean-IoU.

researchpythongo
0
3
Vggsounder EvalA

Evaluates audio-visual foundation and embedding models on multi-label video classification, probing their ability to recognize sound and visual events across different input modalities. It specifically measures modality alignment, unimodal versus multimodal performance, and susceptibility to distraction from irrelevant background audio or static visuals. Use when the user wants to benchmark on VGGSounder, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Vggsound Sep EvalA

Evaluates zero-shot language-queried audio source separation on human actions, sound-emitting objects, and human-object interactions. The benchmark tests isolation of a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on VGGSound, or asks about evaluating this task. Reports SDRi.

researchpythongo
0
3
Vggsound Continual EvalA

Evaluates a model's ability to perform continual audio-visual classification across sequential tasks without catastrophic forgetting. It measures how well the model retains performance on previously learned categories while learning new ones, across audio, visual, and cross-modal fusion settings. Use when the user wants to benchmark on VGGSound-Instruments, VGGSound-100, VGG-Sound Source, or asks about evaluating this task. Reports Average accuracy.

researchpythongo
0
3
Vga Gui Comprehension EvalA

Evaluates a vision-language model's ability to understand graphical user interfaces (GUIs) and answer user questions based on visual content. It specifically probes the model's capacity to avoid hallucinations by grounding responses in actual GUI elements rather than relying solely on textual priors. Use when the user wants to benchmark on GUI Comprehension Bench, or asks about evaluating this task. Reports GPT evaluation score.

researchpythongo
0
3
Vga Bench EvalA

Evaluates text-to-video generation models across three dimensions: aesthetic quality, aesthetic tagging, and generation fidelity. It measures how well automated neural assessors align with human judgments and ranks models based on normalized scores across multiple visual and formal dimensions. Use when the user wants to benchmark on VGA-Bench, or asks about evaluating this task. Reports five-class accuracy.

researchpythongo
0
3
Vg Cot EvalA

Evaluates the visual reasoning and grounding capabilities of Large Vision-Language Models (LVLMs) by measuring the quality of their step-by-step rationales, the accuracy of their final answers, and the alignment between the generated reasoning and the prediction. Use when the user wants to benchmark on VG-CoT, or asks about evaluating this task. Reports Rationale Quality (RQ), Answer Accuracy (AA), Reasoning-Answer Alignment (RAA).

researchpythonrust
0
3
Vevo Voice Imitation EvalA

This benchmark evaluates a model's ability to perform zero-shot voice imitation by disentangling linguistic content, speaker timbre, and vocal style (accent/emotion). It probes the model's capacity to generate high-intelligibility speech that accurately transfers the target speaker's identity and stylistic attributes from a reference clip without task-specific fine-tuning. Use when the user wants to benchmark on Vevo Evaluation Set (AB, CV, ACCENT, EMOTION), or asks about evaluating this task...

researchpythongo
0
3
Veu Bench EvalA

VEU-Bench evaluates a model's ability to understand video editing by probing 19 fine-grained tasks across 10 dimensions (e.g., shot size, cut types, transitions) and three cognitive stages: recognition, reasoning, and judging. It tests whether models can identify editing components, infer their functions, and judge their effects in video content. Use when the user wants to benchmark on VEU-Bench, or asks about evaluating this task. Reports Scoreall.

researchpythongo
0
3
Versebench EvalA

Evaluates a unified multimodal model's ability to generate synchronized audio and video from text, phonemes, and reference media. It probes zero-shot voice cloning fidelity, lip-sync accuracy, acoustic quality, and cross-modal temporal alignment. Use when the user wants to benchmark on VerseBench, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Veroeval EvalA

Evaluates general visual reasoning capabilities across a diverse set of 30 benchmarks spanning six task categories, including chart/OCR, STEM, spatial/action, knowledge/recognition, grounding, and captioning/instruction following. Use when the user wants to benchmark on VeroEval, or asks about evaluating this task. Reports overall averages.

researchpythongo
0
3
Verite EvalA

Evaluates multimodal misinformation detection models on real-world and synthetic image-caption pairs, specifically probing their ability to distinguish truthful content from out-of-context (OOC) and miscaptioned (MC) misinformation while measuring susceptibility to unimodal bias. Use when the user wants to benchmark on VERITE, COSMOS, VMU-Twitter, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Verisoftbench EvalA

This benchmark probes an AI system's ability to perform repository-scale formal verification in Lean 4. It specifically tests context-aware proof automation, measuring how well models handle project-specific abstractions and transitive dependency closures beyond standard mathematical libraries. Use when the user wants to benchmark on VeriSoftBench-Full, or asks about evaluating this task. Reports solve_rate.

researchpythongo
0
3
Verifact EvalA

Evaluates the factual correctness of long-form LLM-generated responses by decomposing them into atomic facts, detecting and refining incomplete or missing information, and verifying each fact against external web evidence. Use when the user wants to benchmark on Long-form LLM responses, or asks about evaluating this task. Reports Supported/Contradicted/Undecided classification accuracy.

researchpythongo
0
3
Vera EvalA

Evaluates the reasoning capabilities of voice and multimodal models under real-time streaming constraints, quantifying the performance gap between text and voice modalities on tasks with well-defined ground truth. Use when the user wants to benchmark on VERA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Venusbench Gd EvalA

This benchmark evaluates GUI grounding capabilities across a hierarchical taxonomy of basic (element, visual, spatial) and advanced (functional, reasoning, refusal) tasks. It probes a model's ability to accurately locate UI elements in screenshots and handle complex, domain-specific, or unanswerable instructions across web, mobile, and desktop platforms. Use when the user wants to benchmark on VenusBench-GD, or asks about evaluating this task. Reports accuracy.

researchpython
0
3
Velocity Dealiasing EvalA

Evaluates a U-Net model's ability to predict velocity fold numbers and produce dealiased radar velocity fields from folded inputs. It measures both classification accuracy for fold detection and reconstruction fidelity via velocity error metrics. Use when the user wants to benchmark on WSR-88D Level-II/III Radar Data, or asks about evaluating this task. Reports velocity RMSE.

researchpythongo
0
3