Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,870
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,0733,096 of 22,870 skills

Vtab Md Fewshot EvalA

Evaluates few-shot classification performance across diverse visual domains by comparing transfer learning and meta-learning approaches on a unified benchmark combining VTAB and Meta-Dataset. Use when the user wants to benchmark on VTAB+MD, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Vsi Bench EvalA

Evaluates multimodal large language models' ability to reason about 3D spatial relationships from visual inputs. It probes capabilities across configurational tasks (counting, relative distance/direction, route planning), measurement estimation (object/room size, absolute distance), and spatiotemporal ordering. Use when the user wants to benchmark on VSI-Bench, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Vsdx EvalA

Evaluates vision-language models' ability to perceive and reason about non-RGB sensor data (thermal, depth, X-ray). It probes low-level perception (existence, counting, position, description) and high-level understanding (contextual reasoning, sensor-specific physical property interpretation). Use when the user wants to benchmark on VS-TDX, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vr Bench EvalA

This benchmark evaluates the spatial reasoning and trajectory planning capabilities of video generation models and vision-language models through maze-solving tasks. It probes whether models can generate coherent, rule-compliant movement sequences or videos that faithfully navigate complex, multi-type mazes such as regular, irregular, 3D, Sokoban, and trap fields. Use when the user wants to benchmark on VR-Bench, or asks about evaluating this task. Reports MF.

researchpythongo
0
3
Vqa Generalization EvalA

Evaluates a model's ability to answer visual questions by generalizing from synthetic template-based training data to complex, human-written questions. It probes both closed-form accuracy and open-form reasoning capabilities across 3D-rendered and medical imaging domains. Use when the user wants to benchmark on CLEVR-Human, VQA-RAD, SLAKE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vqa Gen EvalA

This benchmark evaluates a model's ability to generalize in Visual Question Answering under coordinated visual and textual distribution shifts. It probes robustness to image corruptions, style transfers, and linguistic variations by measuring in-domain and cross-domain accuracy. Use when the user wants to benchmark on VQA-GEN, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Vqa Cot Reasoning EvalA

Probes the ability of vision-language models to perform multi-step chain-of-thought reasoning on visual inputs across diverse domains like charts, documents, science diagrams, and math. It measures both direct answer accuracy and structured reasoning accuracy. Use when the user wants to benchmark on A-OKVQA, ChartQA, DocVQA, InfoVQA, TextVQA, AI2D, ScienceQA, MathVista, OCRBench, MMStar, MMMU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vqa Captioning EvalA

Evaluates multimodal language models on visual question answering and image captioning tasks, probing their zero-shot and few-shot in-context learning capabilities with interleaved image-text inputs. Use when the user wants to benchmark on OKVQA, TextVQA, COCO, Flickr30k, VQAv2, VizWiz, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vqa Biomedical EvalA

Evaluates multimodal models on biomedical visual question answering tasks. It probes the model's ability to interpret medical images (e.g., X-rays, pathology slides) and generate accurate answers to clinical or radiological questions in both open-ended and closed-ended formats. Use when the user wants to benchmark on VQA-RAD, SLAKE, PathVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vqa Benchmarks EvalA

Evaluates vision-language models' ability to answer questions across general knowledge, OCR, mathematics, and science domains using few-shot in-context learning. It also probes the model's capacity to attend to interleaved image-text contexts through a 'cheat test' protocol. Use when the user wants to benchmark on TextVQA, OKVQA, MathVista, MathVision, MathVerse, ScienceQA-IMG, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vpt Minecraft EvalA

Evaluates an agent's ability to perform complex, multi-step sequential decision-making tasks in a 3D sandbox environment (Minecraft) using a native human-like interface. It probes zero-shot generalization, behavioral cloning fine-tuning, and reinforcement learning fine-tuning for long-horizon crafting and exploration. Use when the user wants to benchmark on webClean, contractor_house, earlygame_keyword, or asks about evaluating this task. Reports reliability.

researchpythonperformance
0
3
Vprochart EvalA

Evaluates a model's ability to understand chart visuals and perform multi-step numerical and logical reasoning to answer natural language questions. It specifically probes visual perception alignment and programmatic solution reasoning over structured chart data. Use when the user wants to benchmark on ChartQA, PlotQA, DVQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Voxtream2 EvalA

This evaluation probes a full-stream text-to-speech model's ability to generate intelligible, natural-sounding speech while dynamically controlling the speaking rate in real-time. It measures objective intelligibility, speaker similarity, audio quality, generation latency, and the accuracy of speaking-rate control against target rates. Use when the user wants to benchmark on Emilia speaking-rate dataset, or asks about evaluating this task. Reports WER (%).

researchpythonperformance
0
3
Voxstream Tts EvalA

Evaluates a streaming text-to-speech model's ability to generate high-quality, speaker-similar, and intelligible audio from text prompts in both non-streaming and full-stream (word-by-word) settings. It probes zero-shot voice cloning, cross-sentence continuity, and real-time synthesis latency. Use when the user wants to benchmark on LibriSpeech test-clean, SEED-TTS test-en, or asks about evaluating this task. Reports Naturalness.

researchpythongo
0
3
Voxprivacy EvalA

Evaluates how well speech language models handle interactional privacy across three tiers: obeying direct secrecy commands, using voice identity for conditional access, and proactively inferring contextually sensitive information to withhold secrets. Use when the user wants to benchmark on VoxPrivacy Benchmark, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Voxpopuli EvalA

Benchmarks multilingual speech representation learning and semi-supervised ASR/ST performance across multiple languages and domains, measuring phoneme discriminability, recognition accuracy, and translation quality. Use when the user wants to benchmark on VoxPopuli, Common Voice, ZeroSpeech 2017, EuroParl-ST, CoVoST 2, or asks about evaluating this task. Reports WER.

researchpythongo
0
3
Voxknesset Demographic EvalA

Evaluates whether the VoxKnesset dataset encodes meaningful demographic signals (gender, religion, birthplace) in pretrained speech representations. It probes the dataset's metadata quality and demographic coverage by training lightweight classifiers on extracted embeddings. Use when the user wants to benchmark on VoxKnesset Speaker-Attributed Longitudinal Subset, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Voxceleb1 Verification EvalA

Evaluates speaker verification capability by measuring how well a model distinguishes same-speaker from different-speaker audio pairs. It probes robustness across different evaluation conditions, including a standard test set, an extended large-scale set, and a constrained same-nationality/gender set. Use when the user wants to benchmark on VoxCeleb1, or asks about evaluating this task. Reports EER (%).

researchpythonperformance
0
3
Vox Safe Bench EvalA

Evaluates social alignment in speech language models across safety, fairness, and privacy dimensions. It distinguishes between content-centric risks (Tier 1) where text alone suffices to trigger norms, and audio-conditioned risks (Tier 2) where benign transcripts become unsafe due to speaker identity, paralinguistic cues, or environmental context. Use when the user wants to benchmark on VoxSafeBench, or asks about evaluating this task. Reports RtA.

researchpythongo
0
3
Votreuth Rehab Depth EvalA

Evaluates marker-less 2D pose estimation models on rehabilitation-specific depth images. It probes the model's ability to generalize from generic adult standing poses to complex clinical postures, including children, and tests robustness against varying subject scales and positions. Use when the user wants to benchmark on ITOP, VtR, VtR-O, or asks about evaluating this task. Reports PCK.

researchpythonperformance
0
3
Vos Language ReferringA

Evaluates a model's ability to perform pixel-level video object segmentation guided by natural language referring expressions, testing both language grounding and temporal consistency in dynamic scenes. Use when the user wants to benchmark on DAVIS-16, DAVIS-17, or asks about evaluating this task. Reports performance score.

researchpythonexpress
0
3
Voldor EvalA

Evaluates monocular visual odometry accuracy and depth estimation quality on urban/highway driving sequences and indoor environments. Probes robustness to non-Gaussian optical flow noise and scale ambiguity without relying on hand-crafted features or loop closure. Use when the user wants to benchmark on KITTI odometry benchmark, KITTI stereo benchmark, TUM RGB-D dataset, or asks about evaluating this task. Reports Trans. error (%), Rot. error (deg/m).

researchpythonperformance
0
3
Void EvalA

Evaluates a video generation model's ability to remove specified objects and their downstream physical interactions (e.g., collisions, shadows, reflections) while maintaining temporal consistency and visual quality. It probes counterfactual reasoning and intuitive physics simulation in dynamic scenes. Use when the user wants to benchmark on Real-world object removal dataset, Synthetic counterfactual dataset, or asks about evaluating this task. Reports Win %.

researchpython
0
3
Voiceloop Tts EvalA

Evaluates a text-to-speech model's ability to synthesize perceptually natural speech and accurately mimic speaker identities from text and reference embeddings. It measures robustness across clean benchmarks, multi-speaker corpora, and noisy in-the-wild recordings, while testing few-shot voice fitting capabilities. Use when the user wants to benchmark on LJ (LJSpeech), Nancy (Blizzard 2011), Blizzard 2013 Audiobook, VCTK, In-the-wild YouTube speeches, or asks about evaluating this task. Repor...

researchpythongo
0
3