Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill flores-200-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Flores 200 Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-flores-200-eval)More formats (shields.io, HTML) on the badges page.
---
name: flores-200-eval
description: Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.
metadata:
skill_kind: dataset_eval
source_arxiv: 2207.04672
bibtex_key: nllb2022scaling
confidence: high
---
# flores-200-eval
> No Language Left Behind: Scaling Human-Centered Machine Translation — NLLB Team et al. (2022) (arXiv:2207.04672, 2022)
## What this evaluates
Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity.
## Datasets
- **FLORES-200** — total 1000; splits: test (1000); repo https://github.com/facebookresearch/fairseq/tree/nllb
## Metrics
- `spBLEU` — range: [0, 1]
- BLEU score using a standardized SentencePiece tokenizer covering 101+ languages, combining 1-gram to 4-gram precision with a brevity penalty.
- `chrF++` — range: [0, 1]
- Character n-gram F-score (n=1 to 6) complemented with word unigrams and bi-grams, computed using sacrebleu settings.
- `XSTS` **(primary)** — range: [1, 5]
- Crosslingual Semantic Text Similarity score on a 1-5 scale assessing meaning preservation. Calibrated scores are computed by taking the median of 3 annotator scores per sentence, averaging across sentences, and applying moderated calibration to adjust for annotator harshness/generosity.
## Input / output format
**Input**: Source sentence and machine-translated target sentence (for human evaluation); source sentence and reference translation (for automatic metrics).
**Output**: Human: integer score from 1 to 5. Automatic: continuous float score.
## Scoring recipe
```python
# 1. Aggregate raw human scores
median_scores = [median(annotator_scores) for annotator_scores in sentence_scores]
H_raw = mean(median_scores)
# 2. Compute calibration offset
C_raw = mean([median(annotator_scores) for annotator_scores in calibration_scores])
C_mean = 3.01 # mean across all directions
# 3. Apply moderated calibration
C_shift = C_raw - C_mean
S = tanh(-C_shift)
E = tanh(H_raw - 5) if C_shift <= 0 else tanh(H_raw - 1)
H_cal = H_raw + S * E
```
## Common pitfalls
- Using uncalibrated human scores, which vary significantly due to annotator harshness/generosity across language pairs.
- Assuming BLEU is sufficient for non-space-separated languages without using spBLEU or chrF++.
- Confusing corpus-level chrF++ with average sentence-level chrF++, which show different correlation strengths with human judgment.
## Evidence (verbatim from paper)
> We find that automated metrics like spBLEU and chrF++ correlate reasonably well with calibrated human evaluations of translation quality, as seen in Figure 24. In particular, we find that the Spearman R correlation coefficients between aggregated XSTS and spBLEU, chrF++ (corpus) and chrF++ (average sentence-level) are 0.710, 0.687, and 0.694 respectively.
## Citation
```bibtex
@misc{nllb2022scaling,
title={No Language Left Behind: Scaling Human-Centered Machine Translation},
author={NLLB Team et al. (2022)},
year={2022},
note={arXiv:2207.04672}
}
```
- arXiv: 2207.04672
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!