Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,870
skills in category
953
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 3,0013,024 of 22,870 skills

Wildbe EvalA

Evaluates object detection algorithms on drone-captured images of wild berries in cluttered, dynamic forest environments. It probes localization and classification capabilities under severe lighting variations, occlusion, and cross-domain transfer settings (different areas, cameras, and datasets). Use when the user wants to benchmark on WildBe, or asks about evaluating this task. Reports Average Precision (AP).

researchpythongo
0
3
Wildasr EvalA

This benchmark probes the robustness of automatic speech recognition (ASR) systems under realistic, out-of-distribution conditions. It specifically evaluates performance degradation across environmental noise, demographic shifts (accent, age, child speech), and linguistic diversity (short, incomplete, code-switched utterances), while also measuring semantic hallucination rates beyond standard lexical error metrics. Use when the user wants to benchmark on WildASR, or asks about evaluating this...

researchpythongit
0
3
Wild Tab EvalA

Evaluates the out-of-distribution (OOD) generalization capability of tabular regression models by measuring performance gaps between in-distribution and out-of-distribution test sets. It probes whether advanced OOD training strategies or complex architectures can reliably outperform simple Empirical Risk Minimization (ERM) on unseen data distributions. Use when the user wants to benchmark on VPower_S, VPower_R, Weather, or asks about evaluating this task. Reports MAE.

researchpythonexpress
0
3
Wikitq EvalA

Evaluates a model's ability to perform complex tabular reasoning to answer open-ended questions based on a provided table. It probes the model's capacity to extract, aggregate, and filter information from structured data to produce short text span answers. Use when the user wants to benchmark on WikiTQ, or asks about evaluating this task. Reports denotation accuracy.

researchpythongo
0
3
Wikitablequestions EvalA

Evaluates a model's ability to perform semantic parsing over semi-structured tabular data by generating database queries or answers from natural language questions. It measures how well the model aligns textual utterances with table schemas and content to retrieve correct results. Use when the user wants to benchmark on WIKITABLEQUESTIONS, or asks about evaluating this task. Reports execution accuracy.

researchpythongo
0
3
Wikisql EvalA

Evaluates a model's ability to translate natural language questions into executable SQL queries against a given database schema. It probes semantic parsing, table understanding, and the generation of syntactically and semantically correct structured queries. Use when the user wants to benchmark on WikiSQL, or asks about evaluating this task. Reports Acc_ex.

researchpythongo
0
3
Wikipedia Vandal Detection EvalA

Evaluates a model's ability to detect malicious Wikipedia editors (vandals) using only benign user data for training. It probes one-class anomaly detection and sequential behavior modeling by measuring how well the system distinguishes benign from malicious users based on edit sequences. Use when the user wants to benchmark on UMDWikipedia, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Wikimatrix EvalA

Assesses the quality of automatically mined parallel sentence pairs by training neural machine translation models and measuring their downstream translation accuracy. Use when the user wants to benchmark on WikiMatrix, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Wikilingua EvalA

This benchmark evaluates cross-lingual abstractive summarization, specifically the ability of models to generate coherent English summaries from articles written in other languages. It probes how well systems can handle translation and summarization jointly, either through direct cross-lingual fine-tuning or two-step pipeline approaches. Use when the user wants to benchmark on WikiLingua, or asks about evaluating this task. Reports ROUGE-L F1.

researchpythongo
0
3
Wikihow EvalA

Evaluates text summarization systems on procedural, step-by-step articles written by non-journalists. It probes the model's ability to handle long sequences, non-inverted-pyramid structures, and high-abstraction content compared to standard news datasets. Use when the user wants to benchmark on WikiHow, or asks about evaluating this task. Reports ROUGE-L.

researchpythongo
0
3
Wikidata Ned EvalA

Evaluates a model's ability to disambiguate named entities in text by matching them to correct Wikidata entries using graph-based representations. It probes how well different neural architectures leverage graph triplet information versus full graph topology for entity resolution. Use when the user wants to benchmark on Wikidata-Disamb, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Wikichat Simulated Dialogue EvalA

Evaluates the factual accuracy, conversational quality, and latency of knowledge-grounded chatbots in simulated multi-turn dialogues across head, tail, and recent knowledge domains. Use when the user wants to benchmark on Simulated Dialogues (WikiChat), or asks about evaluating this task. Reports factual_accuracy.

researchpythongit
0
3
Wikicatsum Rouge EvalA

Evaluates abstractive multi-document summarization models on their ability to generate coherent, content-adequate summaries across three domains (Company, Film, Animal). The protocol measures lexical and sentence-level overlap between generated summaries and reference summaries using ROUGE metrics, while also contextualizing scores against a baseline overlap between input documents and summaries. Use when the user wants to benchmark on WIKICATSUM, or asks about evaluating this task. Reports R...

researchpythonperformance
0
3
Wikiasp EvalA

Evaluates multi-domain aspect-based summarization, requiring models to first discover relevant aspects (Wikipedia section titles) from cited references and then generate domain-specific summaries. It probes content selection, cross-document pronoun resolution, and temporal ordering in multi-source generation. Use when the user wants to benchmark on WikiAsp, or asks about evaluating this task. Reports R-2.

researchpythongo
0
3
Wiki Eval EvalA

Measures how well automated RAG scoring metrics align with human preferences in pairwise comparison tasks. It probes the ability of reference-free faithfulness, answer relevance, and context relevance estimators to replicate human judgment on answer and context quality. Use when the user wants to benchmark on WikiEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Wiki Alumni EvalA

Evaluates knowledge graph models on node and entity classification under edge incompleteness, and link prediction with varying feature/class configurations. It probes how well models leverage relational structure, node features, and joint link prediction capacity to handle missing edges and predict missing links. Use when the user wants to benchmark on WikiAlumni, or asks about evaluating this task. Reports validation accuracy.

researchpythongo
0
3
Widspeech Bench EvalA

Evaluates end-to-end speech-to-speech (S2S) language models on real-world conversational tasks, probing their ability to handle diverse query types, paralinguistic features (prosody, disfluencies), and robustness to background noise. Use when the user wants to benchmark on WildSpeech-Bench, or asks about evaluating this task. Reports Score.

researchpythongo
0
3
Wider Face EvalA

Evaluates face detection algorithms on real-world images with extreme variations in scale, pose, occlusion, and event context. It probes the ability of detectors to handle small faces, heavy occlusion, and atypical poses under standard bounding box matching criteria. Use when the user wants to benchmark on WIDER FACE, or asks about evaluating this task. Reports Average Precision (AP).

researchpythongo
0
3
Wide Search EvalA

Evaluates a model's ability to perform broad information seeking by decomposing complex queries into parallel subtasks and producing structured tabular outputs. It also measures robustness on standard single-hop and multi-hop open-domain QA tasks. Use when the user wants to benchmark on WideSearch, or asks about evaluating this task. Reports item F1 score.

researchpythongo
0
3
Wide Deep Recommender EvalA

Evaluates a hybrid recommender model's ability to balance memorization of frequent user-item interactions with generalization to unseen combinations for ranking candidate apps. The protocol measures predictive accuracy on a static holdout set and business impact via live A/B testing on user acquisition rates. Use when the user wants to benchmark on Google Play App Store (internal), or asks about evaluating this task. Reports Online Acquisition Gain.

researchpythongo
0
3
Whisper Zero Shot EvalA

Evaluates the zero-shot generalization capability of a speech recognition model across diverse English and multilingual domains. It measures robustness to out-of-distribution audio, varying noise levels, and translation tasks without any dataset-specific fine-tuning. Use when the user wants to benchmark on LibriSpeech, Common Voice, Fleurs, CoVoST2, Multilingual LibriSpeech (MLS), VoxPopuli, or asks about evaluating this task. Reports WER.

researchpythonperformance
0
3
Whamr EvalA

Evaluates single-channel speech separation and enhancement (denoising/dereverberation) capabilities under realistic noisy and reverberant conditions using synthetically generated reverberant mixtures. Use when the user wants to benchmark on WHAMR!, or asks about evaluating this task. Reports SI-SDR.

researchpython
0
3
Wep Verbalization Validity EvalA

Evaluates neural language models' ability to understand Words of Estimative Probability (WEP) by testing their capacity to distinguish valid from invalid probabilistic verbalizations and perform logical consistency checks in probabilistic reasoning. Use when the user wants to benchmark on WEP Reasoning 1 hop, WEP Reasoning 2 hops, WEP-UNLI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Wenetspeech Wu Bench EvalA

Evaluates speech processing capabilities for the Chinese Wu dialect, including automatic speech recognition (ASR), automatic speech translation (AST), speaker attribute prediction (gender, age), emotion recognition, text-to-speech (TTS), and instruction-following TTS. Use when the user wants to benchmark on WenetSpeech-Wu-Bench, or asks about evaluating this task. Reports CER (%).

researchpythongo
0
3