Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 2,809–2,832 of 22,846 skills
This benchmark evaluates a model's ability to perform word-level quality estimation across multiple language pairs, identifying whether translated words are correct ('OK') or incorrect ('BAD'), as well as detecting target gaps and source-side error triggers. It probes cross-lingual transfer and fine-grained alignment-aware error detection in machine translation. Use when the user wants to benchmark on WMT QE datasets (En-Zh, En-Cs, En-De, En-Ru, En-Lv, De-En), or asks about evaluating this ta...
Evaluates how well multi-modal models learn word meanings and semantic relationships from limited data, comparing visual grounding against language-only baselines across word-relatedness, feature prediction, and POS tagging tasks. Use when the user wants to benchmark on SimLex-999, SimVerb-3500, Word-relatedness dataset (Bruni et al., 2012), or asks about evaluating this task. Reports human-likeness measure.
Evaluates the quality of multilingual word embeddings by measuring semantic similarity/relatedness and word categorization accuracy. It probes whether visual grounding improves cross-lingual semantic alignment and clustering of basic-level concepts. Use when the user wants to benchmark on WordSim353, MEN, RW, MTurk, simVerb, SimLex999, Battig, AP, BLESS, ESSLLI-a, ESSLLI-b, ESSLLI-c, Almarsoomi, MC30, Saif40, WordSim, or asks about evaluating this task. Reports Spearman correlation.
Evaluates the ability of LLMs to accurately assess machine translation quality across varying input lengths (segment, document, and long-form). It probes whether LLMs can maintain consistent error detection and system ranking accuracy when processing longer texts, and tests prompting/fine-tuning strategies to mitigate length bias. Use when the user wants to benchmark on WMT'24 metrics shared task, or asks about evaluating this task. Reports system-level pairwise accuracy.
Evaluates machine translation quality for low-resource Northeast Indian languages (Assamese, Khasi, Mizo, Manipuri) paired with English. It probes cross-lingual transfer capabilities, model adaptation under data scarcity, and the effectiveness of architectural constraints like layer freezing and script-based language grouping. Use when the user wants to benchmark on IndicNECorp1.0, or asks about evaluating this task. Reports BLEU.
Evaluates document-level machine translation quality using multi-turn conversational prompting strategies with LLMs, measuring contextual coherence and translation accuracy across multiple language directions and domains. Use when the user wants to benchmark on WMT 24 General Track, WMT 23 Chinese-to-English, or asks about evaluating this task. Reports dBLEU.
Evaluates machine translation systems for bilingual customer support conversations, focusing on context utilization, discourse coherence, and turn-level versus conversation-level translation quality across five language pairs. Use when the user wants to benchmark on MAIA 2.0, or asks about evaluating this task. Reports COMET.
Evaluates machine translation systems across 55 languages and dialects using automatic metrics and significance testing to compare translation quality across different domains and language pairs. Use when the user wants to benchmark on WMT24++, or asks about evaluating this task. Reports BLEU.
Evaluates machine translation models on English-to-multiple-target-language pairs using standard development sets. It measures translation quality via BLEU scores to compare bilingual versus multilingual decoder representations and capacity. Use when the user wants to benchmark on WMT22 General Machine Translation, Multitarget TED talks, or asks about evaluating this task. Reports BLEU.
Evaluates sign language translation from video to spoken text. It probes the model's ability to handle long videos, large vocabularies, and high singleton rates by leveraging full-body and lip-reading visual features. Use when the user wants to benchmark on WMT 2022 Shared Task, PHOENIX 2014T, or asks about evaluating this task. Reports BLEU.
Evaluates machine translation quality estimation systems by predicting human judgments on translation adequacy and fluency (Direct Assessment) and classifying translation errors (CED). It probes the model's ability to correlate predicted scores with human ratings and accurately detect translation quality issues across multiple language pairs. Use when the user wants to benchmark on WMT 2021 Quality Estimation Shared Task datasets, or asks about evaluating this task. Reports Pearson's correlat...
Evaluates neural machine translation performance across news and biomedical domains for English-German and English-Russian language pairs. It probes the model's ability to handle domain-specific vocabulary, cross-lingual alignment, and translation quality under constrained data conditions typical of shared task tracks. Use when the user wants to benchmark on WMT21 News & Biomedical Shared Tasks, or asks about evaluating this task. Reports BLEU.
Evaluates the translation quality and inference speed of non-autoregressive versus autoregressive machine translation models on English-German news text. It probes the practical trade-offs between decoding latency and translation accuracy under realistic deployment conditions. Use when the user wants to benchmark on WMT21 News Translation, or asks about evaluating this task. Reports BLEU.
Evaluates Chinese-to-English machine translation performance on biomedical texts. It measures how well a neural MT system can translate domain-specific terminology and syntax while handling case, punctuation, and subword tokenization conventions. Use when the user wants to benchmark on WMT21 OK-aligned biomedical test set, or asks about evaluating this task. Reports BLEU.
Evaluates document-level and literary-domain machine translation quality, focusing on discourse-aware metrics like consistency, anaphora resolution, and term consistency across long-form Chinese web novels. The protocol combines automated n-gram and neural metrics with a structured human judgment framework to capture cross-sentence coherence and literary style. Use when the user wants to benchmark on WMT 2024 Discourse-Level Literary Translation Shared Task, or asks about evaluating this task...
Evaluates machine translation quality estimation across sentence-level scoring, word-level error detection, and fine-grained error span classification. It probes a model's ability to predict translation quality and localize specific errors without relying on human references during inference. Use when the user wants to benchmark on WMT2022 QE EN-DE dataset, WMT2022 Metric EN-DE dataset, WMT17/19/20 Post-editing EN-DE datasets, or asks about evaluating this task. Reports MCC.
Evaluates machine translation systems on discourse-level literary translation from Chinese to English, focusing on long-range context, cultural adaptation, and stylistic fidelity across entire web novels rather than isolated sentences. Use when the user wants to benchmark on WMT 2023 Discourse-Level Literary Translation Test Set, or asks about evaluating this task. Reports d-BLEU.
Evaluates neural machine translation systems across multiple language pairs (EN-DE, EN-CS, CS-EN, EN-RO, RO-EN, EN-RU, RU-EN) on news text. It measures translation quality using BLEU scores on held-out test sets to assess the impact of techniques like back-translation, ensembling, and subword segmentation. Use when the user wants to benchmark on WMT 2016 News Translation, or asks about evaluating this task. Reports BLEU.
Evaluates Automatic Post-Editing (APE) systems by measuring how well they correct machine-translated German sentences using monolingual and bilingual neural translation models combined via log-linear weighting. Use when the user wants to benchmark on WMT 2016 APE Shared Task Development Set, or asks about evaluating this task. Reports TER.
Evaluates machine translation quality by scoring system-generated English translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 ZH-EN, or asks about evaluating this task. Reports COMET.
Evaluates machine translation quality by scoring system-generated German translations against human reference sentences and expert MQM scores. It measures how well automatic metrics correlate with human judgments across different translation systems. Use when the user wants to benchmark on WMT20 EN-DE, or asks about evaluating this task. Reports COMET.
Evaluates machine translation quality between similar languages (Czech to Polish) using a multi-encoder transformer trained on out-of-domain data filtered by cross-entropy differences. Probes the model's ability to adapt to low-resource similar language pairs via domain adaptation and data selection. Use when the user wants to benchmark on WMT19 SLT Shared Task dataset, or asks about evaluating this task. Reports BLEU.
Evaluates an online learning framework's ability to dynamically identify the highest-quality machine translation systems from an ensemble using minimal human feedback, and measures the sample efficiency (number of human assessments needed) to converge to the official top-performing systems. Use when the user wants to benchmark on WMT'19 News Translation, or asks about evaluating this task. Reports convergence_to_top3.
Evaluates the effectiveness of automatic web data selection for domain-specific machine translation in the news domain. It probes whether a document-level topic classifier can filter noisy parallel data to improve MT system performance compared to state-of-the-art baselines on the WMT-18 benchmark. Use when the user wants to benchmark on WMT-18 News Shared Task, or asks about evaluating this task. Reports BLEU.