Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,860
skills in category
995
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,169–6,192 of 23,860 skills

Multivox EvalA

Evaluates voice assistants' ability to jointly ground visual and paralinguistic speech cues (e.g., pitch, emotion, volume, background sounds) in context-aware responses. It specifically tests robustness against confounding samples that flip speech properties to prevent overreliance on unimodal priors. Use when the user wants to benchmark on MultiVox, or asks about evaluating this task. Reports visual grounding and non-verbal speech signals.

researchpythonperformance
0
3
Multiview Cir EvalA

This benchmark evaluates a model's ability to perform product-level composed image retrieval (CIR) in fashion e-commerce, specifically handling multi-view product images and short modification queries. It probes the model's capacity to align visual perception with textual reasoning across multiple views while filtering out irrelevant gallery items. Use when the user wants to benchmark on DeepFashion, Fashion200K, FashionGen-val, or asks about evaluating this task. Reports Recall@5.

researchpythongo
0
3
Multiverse EvalA

Evaluates the multi-turn conversational reasoning and sustained dialogue capabilities of Vision-Language Models (VLMs) across diverse domains like mathematics, coding, and creative tasks. It probes how well models leverage dialogue history (in-context learning) and maintain consistency over extended interactions. Use when the user wants to benchmark on MultiVerse, or asks about evaluating this task. Reports checklist-based evaluation.

researchpythongo
0
3
Multivent2.0 EvalA

Evaluates event-centric video retrieval across six languages, requiring models to match natural language queries about specific world events to relevant long-form videos using multimodal signals (vision, audio, OCR, metadata). It probes a model's ability to integrate cross-lingual, cross-modal information for complex event understanding rather than simple visual matching. Use when the user wants to benchmark on MultiVENT 2.0, or asks about evaluating this task. Reports Retrieval Performance.

researchpythongo
0
3
Multivariate Ts Prediction EvalA

Evaluates the ability of spatiotemporal attention models to accurately predict future values in multivariate time series across environmental, building HVAC, and clinical domains. It also probes the model's capacity to produce interpretable attention weights that align with known physical or physiological relationships. Use when the user wants to benchmark on Beijing PM2.5 Data Set, Building HVAC Dataset, MIMIC-III EHR Dataset, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Multiui EvalA

Evaluates multimodal models' ability to understand and interact with complex webpage UIs, perform text-rich visual grounding, and generalize to OCR and document understanding tasks. Use when the user wants to benchmark on VisualWebBench, Mind2Web, DocVQA, ChartQA, or asks about evaluating this task. Reports element accuracy.

researchpythongo
0
3
Multitask Detox Utility EvalA

Evaluates a fine-tuned LLM's ability to mitigate toxicity while preserving general knowledge and utility across multiple benchmarks. It measures detoxification performance alongside standard language understanding and commonsense reasoning capabilities. Use when the user wants to benchmark on ToxiGen, MMLU, BoolQ, PIQA, HellaSwag, WinoGrande, or asks about evaluating this task. Reports MMLU (utility).

researchpythongo
0
3
Multisports EvalA

Evaluates multi-person spatio-temporal action detection in sports videos, probing the model's ability to localize fine-grained actions across multiple concurrent persons, handle occlusion, and model long-range temporal context. Use when the user wants to benchmark on MultiSports, or asks about evaluating this task. Reports frame-mAP@0.5.

researchpythongit
0
3
Multirobustbench EvalA

Evaluates machine learning model robustness against multiple diverse adversarial attacks (e.g., ℓₚ-norm, color shifts, spatial transformations) across varying strengths. It quantifies how well defenses maintain performance under worst-case and average-case multiattack scenarios, addressing bias from varying attack difficulties. Use when the user wants to benchmark on MultiRobustBench, or asks about evaluating this task. Reports competitiveness ratio (CR).

researchpythonperformance
0
3
Multiref Bench EvalA

Evaluates the ability of image generation models to simultaneously align and incorporate multiple visual reference conditions (e.g., bounding boxes, depth maps, masks, sketches) alongside text instructions. It probes complex multi-source creative synthesis, testing both global image quality and fine-grained reference fidelity across different input formats and processing orders. Use when the user wants to benchmark on MULTIREF-BENCH, or asks about evaluating this task. Reports Overall Assessm...

researchpythontesting
0
3
Multiqt Question Tracking EvalA

Evaluates a model's ability to perform real-time, multimodal sequence labeling to detect and classify questions in emergency call speech. It probes robustness to noisy ASR transcriptions and temporal alignment under streaming conditions. Use when the user wants to benchmark on question and symptoms tracking datasets, or asks about evaluating this task. Reports TIMESTEP F1.

researchpythongo
0
3
Multiq EvalA

Evaluates the multilingual language fidelity and question-answering accuracy of open LLMs across 137 typologically diverse languages. It probes whether models respond in the prompt's language and whether their answers are factually correct, highlighting the impact of tokenization strategies and model scaling on multilingual performance. Use when the user wants to benchmark on MultiQ, or asks about evaluating this task. Reports QA accuracy (%).

researchpythongo
0
3
Multiple Choice Vqa EvalA

This evaluation probes the true multimodal reasoning capability of vision-language models on multiple-choice question answering tasks. It specifically measures whether models rely on actual question understanding or exploit visual relevance imbalances between correct answers and distractors. Performance is assessed under both standard (vision, question, options) and question-omitted (vision, options) settings to detect easy-option bias. Use when the user wants to benchmark on NExT-QA, MMStar,...

researchpythongo
0
3
Multipl E Low Resource EvalA

Evaluates large language models' ability to generate functionally correct code in low-resource programming languages (R and Racket). It probes how well in-context learning strategies and fine-tuning adapt pre-trained models to languages with limited training data and documentation. Use when the user wants to benchmark on MultiPL-E, or asks about evaluating this task. Reports pass@1.

researchpythonperformance
0
3
Multiphysics Parameter Estimation EvalA

Evaluates the accuracy of a joint multiphysics-decision tree learning framework in estimating subsurface transport parameters and simulating state variables (pressure head, temperature, concentration) under stochastic boundary conditions. It benchmarks the reduced-order surrogate models against a full numerical multiphysics inversion baseline. Use when the user wants to benchmark on Stochastic managed aquifer recharge dataset, or asks about evaluating this task. Reports Nash-Sutcliffe efficie...

researchpythongit
0
3
Multiparadetox EvalA

Evaluates text detoxification models across Russian, Ukrainian, and Spanish by measuring how effectively they transform toxic input into neutral output. The benchmark probes a model's ability to remove offensive language while preserving the original semantic content and maintaining grammatical fluency in the target language. Use when the user wants to benchmark on MultiParaDetox, or asks about evaluating this task. Reports STA.

researchpythongo
0
3
Multioff Hateful Meme EvalA

Binary classification of memes as offensive or non-offensive. It probes a model's ability to detect hate speech in multimodal content by leveraging serialized scene graphs and knowledge graph entities alongside raw text. Use when the user wants to benchmark on MultiOFF, or asks about evaluating this task. Reports F1 score (offensive class).

researchpythongo
0
3
Multinrc EvalA

This benchmark evaluates LLMs' ability to perform multi-step reasoning in native non-English languages (French, Spanish, Chinese) across linguistic, wordplay, cultural/tradition, and culturally-grounded math categories. It specifically probes whether models rely on translation bias or possess deep cultural and linguistic contextual knowledge required for accurate problem-solving. Use when the user wants to benchmark on MultiNRC, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodalqa EvalA

Evaluates complex question answering capabilities that require joint reasoning across text, tables, and images. It probes multi-hop reasoning, cross-modal inference, and the ability to align and process structured and unstructured data to produce correct answer lists. Use when the user wants to benchmark on MultiModalQA, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Multimodal Understanding EvalA

Evaluates a model's ability to understand and reason over diverse visual inputs, including general VQA, document/chart understanding, OCR, and hallucination robustness. Use when the user wants to benchmark on MMMU(Val), MMStar, MME, OCRBench, HallB(Avg), MMB(Dev En V1.1), TextVQA, DoCVQA, InfoVQA, AI2D, ChartQA, RWQA, or asks about evaluating this task. Reports VLMEvalKit score.

researchpythongo
0
3
Multimodal Vqa EvalA

Evaluates the zero-shot and few-shot visual question answering capabilities of multimodal large language models (MLLMs). It probes scene and spatial understanding, OCR capabilities, commonsense knowledge reasoning, and multimodal in-context learning across diverse benchmarks. Use when the user wants to benchmark on GQA, VQA-v2, VizWiz, TextVQA, OKVQA, POPE, MMMU (Val), MMBench (Dev), MMStar, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Visual Reasoning EvalA

This evaluation probes a model's ability to perform long-chain, multi-modal reasoning on mathematical problems that require deep visual understanding. It measures how well the model integrates image evidence with textual reasoning steps to arrive at correct answers across diverse math benchmarks. Use when the user wants to benchmark on MathVista, MathVision, MathVerse, Dynamath, OlympiadBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Multimodal Tool Use EvalA

Evaluates the ability of agentic multimodal models to perform visual perception, document understanding, and mathematical reasoning. It specifically probes whether models can strategically decide when to invoke external tools (e.g., image cropping, web search, Python code execution) versus answering directly, balancing task accuracy with tool efficiency. Use when the user wants to benchmark on V-Bench, HRBench-4K/8K, TreeBench, MME-RealWorld, SEEDBench2-Plus, CharXiv, MathVista_mini, MathVers...

researchpythongo
0
3
Multimodal Reward Benchmarks EvalA

Evaluates multimodal reward models on preference ranking tasks across image and video domains, measuring how well they score or rank candidate responses compared to ground-truth preferences. It compares multi-response scoring against single-response baselines and generative judges, while also assessing inference efficiency and downstream policy optimization stability. Use when the user wants to benchmark on VL-RewardBench, Multimodal RewardBench, MM-RLHF RewardBench, MR2Bench-Image, VideoRewa...

researchpythongo
0
3