Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

22,846
skills in category
952
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 2,8332,856 of 22,846 skills

Wmt17 Paraphrase EvalA

Evaluates how well a model's predicted quality scores correlate with human judgments on machine translation output. It probes semantic equivalence and paraphrase detection capabilities in the context of MT evaluation. Use when the user wants to benchmark on WMT17, or asks about evaluating this task. Reports Pearson |r|.

researchpython
0
3
Wmt17 Nmt Benchmark EvalA

Evaluates the training efficiency, inference speed, and translation quality of neural machine translation systems on standard WMT17 benchmarks. It probes the trade-offs between model architecture, hardware acceleration (FP16/INT8), and batching strategies. Use when the user wants to benchmark on WMT17 English-German, WMT17 Russian-English, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Wmt17 Mt EvalA

Evaluates neural machine translation systems across multiple language pairs in news and biomedical domains. Probes translation quality, domain adaptation, and system combination techniques like ensembling and reranking on held-out parallel test sets. Use when the user wants to benchmark on WMT17 News Task, HimL Biomedical Task, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Wmt14 En Fr EvalA

Evaluates the effectiveness of an RNN Encoder-Decoder architecture for statistical machine translation on English-to-French tasks. It measures how well the model learns phrase representations and improves translation quality over a traditional phrase-based baseline system. Use when the user wants to benchmark on WMT'14 English/French, or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Wmt Slt 22 EvalA

Evaluates the capability of sign language to text translation systems on low-resource datasets. It measures how accurately a model can convert 3D pose sequences of sign language into corresponding spoken language text. Use when the user wants to benchmark on FocusNews, SRF, or asks about evaluating this task. Reports BLEU.

researchpythonapi
0
3
Wmt Mt EvalA

This protocol evaluates the machine translation quality of large language models across multiple language pairs. It measures translation accuracy and fluency by comparing model outputs against gold references and state-of-the-art baselines using neural quality estimation metrics. The benchmark probes the model's ability to generalize across diverse language directions and avoid generating near-perfect but flawed translations. Use when the user wants to benchmark on WMT'21 Test Set, WMT'22 Tes...

researchpythongo
0
3
Wmt Bleu EvalA

Evaluates the quality of an unsupervised web-mined parallel corpus by training neural machine translation models on it and measuring translation performance on standard WMT test sets. It probes whether mined pseudo-parallel data can effectively substitute for human-labeled data in both supervised and unsupervised MT training pipelines. Use when the user wants to benchmark on WMT2014 test set, WMT2016 test set, or asks about evaluating this task. Reports BELU.

researchpythonperformance
0
3
Wmt Ape EvalA

Evaluates a model's ability to perform Automatic Post-Editing (APE) by correcting machine-translated German sentences using the original source text. It probes error detection, grammatical correction, and word-copying capabilities in a multi-source sequence-to-sequence setting. Use when the user wants to benchmark on WMT APE, or asks about evaluating this task. Reports case-sensitive BLEU.

researchpython
0
3
Wmdp EvalA

Evaluates large language models' knowledge of hazardous topics in biosecurity, cybersecurity, and chemical security, as well as their general knowledge and fluency. It serves as a proxy for measuring dual-use risk and benchmarking unlearning methods. Use when the user wants to benchmark on WMDP, or asks about evaluating this task. Reports WMDP.

researchpythongo
0
3
Wmabench EvalA

Evaluates Vision-Language Models on atomic world modeling capabilities across perception (spatial, temporal, motion) and prediction (mechanistic simulation, transitive/compositional inference) tasks. It probes whether VLMs possess internal representations of physical causality, dynamics, and multi-step reasoning comparable to human intuition. Use when the user wants to benchmark on WM-ABench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Wizardlm EvalA

Probes instruction-following capability on complex, real-world prompts across diverse domains like coding, math, reasoning, and formatting. It measures how well models handle demanding, multi-step tasks compared to baselines through blind pairwise human comparison. Use when the user wants to benchmark on WizardEval, or asks about evaluating this task. Reports win_rate.

researchpythongit
0
3
Wivi Har EvalA

Evaluates human activity recognition (HAR) performance using only wireless Channel State Information (CSI) signals under varying action segmentation windows (1s, 2s, 3s). It probes the robustness of classification models in privacy-preserving environments where visual data is occluded or unavailable during testing. Use when the user wants to benchmark on WiVi, or asks about evaluating this task. Reports OA.

researchpythontesting
0
3
Wiser EvalA

Evaluates a robot's ability to generalize visuomotor planning to unseen visual signals, referring expressions, and spatial relationships during cube-picking and placing tasks. It probes semantic generalization and robustness to distribution shifts in embodied instruction following. Use when the user wants to benchmark on WISER Benchmark, or asks about evaluating this task. Reports Success.

researchpythongo
0
3
Wire57 EvalA

Evaluates Open Information Extraction systems on their ability to accurately extract relational tuples from text. It probes token-level precision and recall by matching predicted arguments and relations against a fine-grained, manually annotated gold standard. Use when the user wants to benchmark on WiRe57, or asks about evaluating this task. Reports token-weighted F1.

researchpythongo
0
3
Winosemitism EvalA

Evaluates whether language models disproportionately associate harmful stereotypes with marginalized groups (Jewish people or LGBTQ+ subgroups) compared to non-target groups. It also assesses the quality and reliability of automated versus human annotation for constructing community-sourced fairness benchmarks. Use when the user wants to benchmark on WinoSemitism, WinoQueer, or asks about evaluating this task. Reports WinoSem. Score.

researchpythongo
0
3
Winoqueer EvalA

Evaluates anti-LGBTQ+ bias in language models by measuring their tendency to prefer stereotypical completions over counterfactual ones when prompted with identity-specific contexts. Use when the user wants to benchmark on WinoQueer, or asks about evaluating this task. Reports bias score.

researchpythongit
0
3
Winoground EvalA

Probes vision-language models' ability to understand visio-linguistic compositionality and word order sensitivity. The task requires matching images to captions where identical words are rearranged to change the described scene, testing structural grounding rather than lexical overlap. Use when the user wants to benchmark on Winoground, or asks about evaluating this task. Reports image-caption score.

researchpythongo
0
3
Winogender Schemas EvalA

This benchmark probes systematic gender bias in coreference resolution systems by measuring how often models resolve gendered pronouns to occupations differently based solely on pronoun gender. It evaluates whether models reinforce real-world occupational gender disparities and how performance degrades on counter-stereotypical ('gotcha') examples. Use when the user wants to benchmark on Winogender schemas, or asks about evaluating this task. Reports bias_score.

researchpythongo
0
3
WinoBias EvalA

Evaluates gender bias in coreference resolution systems by measuring performance disparity between pro-stereotypical and anti-stereotypical sentences. It probes whether models rely on gender stereotypes when resolving coreferences in challenging, Winograd-style contexts. Use when the user wants to benchmark on WinoBias, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Wingpt 3.0 Benchmark EvalA

Evaluates large language models on comprehensive medical reasoning, clinical calculation, and general cognitive capabilities. It probes domain-specific knowledge application, diagnostic reasoning, and complex problem-solving in real-world clinical and academic settings. Use when the user wants to benchmark on MedCalc, MedReMCQ, CMMLU, MATH-500, MedQA-USMLE, MedMCQA, PubMedQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Window Wise ComplexityA

Quantifies the intrinsic predictability of time series data by measuring window-wise pattern complexity in the frequency domain. It establishes a data-driven performance lower bound for forecasting models and identifies whether standard benchmarks have reached saturation. Use when the user has predictions and gold and needs to compute window-wise complexity.

researchpythongo
0
3
Wind Power Ensemble EvalA

Evaluates the calibration, sharpness, and accuracy of probabilistic wind power forecasts under different ensemble post-processing strategies (raw, weather-only, power-only, and joint weather-power post-processing). It probes whether correcting biases at the weather stage alone is sufficient, or if direct post-processing of the final power ensemble is required to handle non-linear power curve biases. Use when the user wants to benchmark on Benchmark Data, Swedish Data Set, or asks about evalua...

researchpythongit
0
3
Wili 2018 EvalA

Evaluates the ability of models to correctly identify the language of monolingual text paragraphs. It probes language identification capabilities across a wide range of languages (235) with balanced representation. Use when the user wants to benchmark on WiLI-2018, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Wildscore EvalA

This benchmark evaluates multimodal large language models' ability to perform multi-step, context-sensitive reasoning over symbolic musical notation. It probes capabilities in harmonic analysis, rhythmic interpretation, structural form recognition, and expressive markings through multiple-choice questions derived from real-world compositions and forum queries. Use when the user wants to benchmark on WildScore, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3