Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,485
skills in category
979
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,889–3,912 of 23,485 skills

Varex EvalA

Evaluates multi-modal structured data extraction from documents, testing a model's ability to parse visual or textual layouts, adhere to a provided JSON schema, and generate compliant structured outputs. It specifically probes schema compliance, layout understanding, and cross-modal robustness across plain text, spatial text, and image inputs. Use when the user wants to benchmark on VAREX, or asks about evaluating this task. Reports exact match (EM).

researchpythongo
0
3
Vane Bench EvalA

Evaluates the ability of video-language models to detect subtle, rapid, and contextually nuanced anomalies in both real-world surveillance footage and high-fidelity AI-generated videos. It probes fine-grained temporal reasoning, visual grounding, and robustness to synthetic artifacts through a multiple-choice question-answering format. Use when the user wants to benchmark on VANE-Bench, or asks about evaluating this task. Reports MC-Video QA accuracy.

researchpythongo
0
3
Valueground EvalA

Evaluates whether multimodal large language models (MLLMs) can maintain consistent culture-conditioned value judgments when response options are replaced with minimally contrastive visual proxies. It probes cross-modal prediction stability and the ability to ground textual value tendencies in subtle visual contrasts. Use when the user wants to benchmark on ValueGround, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
ValiditysoftA

Evaluates the faithfulness and semantic plausibility of model-agnostic XAI techniques by generating soft counterfactuals via token-level perturbations. It measures whether perturbations actually change model predictions and whether the generated explanations align with the true causal impact of those changes. Use when the user has predictions and gold and needs to compute Validitysoft, Csoft.

researchpythongo
0
3
Valerie22 EvalA

This protocol evaluates the perceptual fidelity and cross-domain generalization capability of the VALERIE22 synthetic urban dataset by training a semantic segmentation model on it and testing on real-world automotive datasets. It specifically probes how dataset diversity (unique 3D assets) and training scale affect downstream perception performance. Use when the user wants to benchmark on VALERIE22, Cityscapes, A2D2, BDD100K, India Driving Dataset, Mapillary Vistas, or asks about evaluating t...

researchpythontesting
0
3
Vaexbench EvalA

Evaluates multimodal large language models' ability to perform extractive and abstractive spatiotemporal reasoning on egocentric videos. It probes long-horizon memory, object tracking, spatial orientation, and metric distance estimation under both multiple-choice and free-form generation settings. Use when the user wants to benchmark on VAEX-Bench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Vae Malware Detection EvalA

Evaluates the effectiveness of Variational Autoencoder (VAE)-derived latent space features for malware classification using traditional machine learning models. It probes robustness to data partitioning, random seed initialization, and computational efficiency without hyperparameter tuning. Use when the user wants to benchmark on EMBER, BODMAS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Vad Prediction EvalA

Evaluates a model's ability to predict continuous emotional dimensions (Valence, Arousal, Dominance) from text. Specifically probes the model's capacity to capture affective polarization signals in parliamentary discourse. Use when the user wants to benchmark on Knesset VAD Annotation, or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Vad Anticipation EvalA

Evaluates a model's ability to detect anomalous events in surveillance videos and anticipate their occurrence in future frames. It specifically probes scene-dependent anomaly recognition and multi-step temporal anticipation. Use when the user wants to benchmark on ShanghaiTech, CUHK Avenue, IITB Corridor, NWPU Campus, ShanghaiTech-sd, or asks about evaluating this task. Reports AUC (%).

researchpythonperformance
0
3
Vabench EvalA

Evaluates the quality, cross-modal consistency, synchronization, and spatial audio rendering of text-to-audio-video and image-to-audio-video generation models. It probes physical plausibility, emotional expressiveness, and stereo separation across seven real-world sound categories. Use when the user wants to benchmark on VABench, or asks about evaluating this task. Reports Audio-Visual Align.

researchpythongo
0
3
Vaani Asr Lid EvalA

This evaluation protocol assesses the utility of the Vaani dataset for fine-tuning automatic speech recognition (ASR) and spoken language identification (LID) models across diverse Indian languages and regions. It measures performance gains from fine-tuning on Vaani's transcribed audio and images against established benchmarks, highlighting regional dialectal variations and low-resource language capabilities. Use when the user wants to benchmark on Vaani, FLEURS, Kathbath, or asks about evalu...

researchpythongit
0
3
V3det EvalA

Probes an object detector's ability to localize and classify instances across a vast, hierarchical vocabulary of over 13,000 categories. It evaluates both closed-set detection performance and open-vocabulary generalization to novel, unseen categories. Use when the user wants to benchmark on V3Det, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
V2x Radar EvalA

Evaluates 3D object detection capabilities for autonomous driving using multi-modal sensors (LiDAR, camera, 4D radar) in single-agent (roadside and vehicle-mounted) and cooperative perception setups. It probes robustness to adverse weather conditions and communication delays in cooperative scenarios. Use when the user wants to benchmark on V2X-Radar, or asks about evaluating this task. Reports AP@IoU.

researchpythongo
0
3
V2v Qa EvalA

Evaluates a multi-modal LLM's ability to fuse 3D perception features from multiple connected vehicles to answer safety-critical driving queries. It probes spatial grounding, notable object identification near planned waypoints, and collision-avoidance trajectory planning in cooperative autonomous driving scenarios. Use when the user wants to benchmark on V2V-QA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
V2v Got Planning EvalA

Evaluates cooperative autonomous driving planning capabilities using a multimodal LLM with graph-of-thoughts reasoning. It measures trajectory prediction accuracy and collision avoidance under occlusion-aware perception and planning-aware prediction scenarios. Use when the user wants to benchmark on V2V-GoT-QA, or asks about evaluating this task. Reports L2 error.

researchpythongo
0
3
V2c Chem Dubbing EvalA

This evaluation probes a model's ability to generate synchronized, emotionally faithful, and speaker-identifiable speech for movie dubbing tasks. It measures audio-visual alignment, spectral similarity, and the preservation of speaker identity and emotional tone against ground-truth recordings. Use when the user wants to benchmark on V2C, Chem, or asks about evaluating this task. Reports LSE-D.

researchpythongo
0
3
V2c Animation EvalA

Evaluates visually-driven voice cloning by measuring speech quality, temporal alignment, length consistency, speaker identity preservation, and emotion transfer accuracy against ground-truth audio. Use when the user wants to benchmark on V2C-Animation, or asks about evaluating this task. Reports MCD-DTW-SL.

researchpythongo
0
3
V Triune EvalA

Evaluates vision-language models on a unified suite of visual reasoning and perception tasks, measuring generalization across real-world benchmarks, mathematical reasoning, and object detection/grounding capabilities. Use when the user wants to benchmark on MEGA-Bench Core, MMMU, MathVista, COCO, OVDEval, CountBench, OCRBench, ScreenSpot-Pro, or asks about evaluating this task. Reports MEGA-Bench Core weighted average.

researchpythongo
0
3
Uvh 26 EvalA

Evaluates object detection models on a domain-specific Indian traffic dataset, probing their ability to localize and classify 14 heterogeneous vehicle types under surveillance viewpoints with varying occlusion and scale. Use when the user wants to benchmark on UVH-26, or asks about evaluating this task. Reports mAP(50:95).

researchpythongo
0
3
Utkface Age Estimation EvalA

This evaluation protocol assesses the accuracy of a lightweight neural network for predicting a person's age from a single facial image. It focuses on regression-based age estimation to determine how well compact models generalize to held-out test data while maintaining deployment efficiency. Use when the user wants to benchmark on UTKFace, or asks about evaluating this task. Reports MAE.

researchpythongit
0
3
Utility Aware Data Pricing EvalA

Evaluates whether token-level quality signals and empirical training gain metrics can accurately predict real data utility for LLM fine-tuning, outperforming traditional row- or token-count baselines. It probes the framework's predictive alignment, ranking fidelity, and robustness to adversarial or low-value data across multiple domains. Use when the user wants to benchmark on Alpaca, GSM8K, CodeXGLUE-Python, or asks about evaluating this task. Reports Spearman rank correlation.

researchpythonperformance
0
3
Utd Mhad EvalA

Evaluates a model's ability to predict future human joint positions over a 15-frame horizon using past observations, while testing continual learning capabilities across different subjects and curriculum-based fine-tuning. Use when the user wants to benchmark on UTD-MHAD, or asks about evaluating this task. Reports MSE.

researchpythontesting
0
3
Usc Asr EvalA

Evaluates automatic speech recognition (ASR) systems on Uzbek language audio by measuring character and word error rates against manually transcribed ground truth. It probes the model's ability to accurately transcribe low-resource speech data without relying on external linguistic resources or pronunciation dictionaries. Use when the user wants to benchmark on USC, or asks about evaluating this task. Reports WER.

researchpythonexpress
0
3
Usb Summarization EvalA

Evaluates multiple text summarization capabilities including extractive/abstractive generation, factuality verification, factual error correction, topic-constrained generation, sentence compression, evidence extraction, and unsupported span detection across diverse domains. Use when the user wants to benchmark on Extractive Summarization (EXT), Abstractive Summarization (ABS), Factuality Classification (FAC), Fixing Factuality (FIX), Topic-based Summarization (TOPIC), Multi-sentence Compressi...

researchpythongo
0
3