Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks. Use when the user wants to benchmark on Multilingual MMLU, TranslatedDolly, Taxi1500, or asks about evaluating this task. Reports accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill muri-101-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Muri 101 Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-muri-101-eval)More formats (shields.io, HTML) on the badges page.
---
name: muri-101-eval
description: Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks. Use when the user wants to benchmark on Multilingual MMLU, TranslatedDolly, Taxi1500, or asks about evaluating this task. Reports accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2409.12958
bibtex_key: koksal2024muri
confidence: high
---
# muri-101-eval
> MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions — Köksal et al. (2024) (arXiv:2409.12958, 2024)
## What this evaluates
Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks.
## Datasets
- **Multilingual MMLU** — total ?; splits: test (-1)
- **TranslatedDolly** — total ?; splits: test (-1)
- **Taxi1500** — total ?; splits: test (-1)
## Metrics
- `accuracy` **(primary)** — range: [0, 1]
- Proportion of correctly predicted labels or answers out of the total number of instances. Standard for classification and multiple-choice benchmarks like MMLU and Taxi1500.
- `win rates` — range: percent
- Percentage of times the evaluated model's generation is preferred over a baseline model's generation when judged by the multilingual Command R+ LLM-as-a-judge.
## Input / output format
**Input**: Instruction prompt with task context, formatted for mT5 instruction tuning. Evaluated with 5-shot or 6-shot in-context examples depending on the benchmark.
**Output**: Free-form text generation completing the instruction.
## Scoring recipe
```python
# For NLU benchmarks (MMLU, Taxi1500)
accuracy = sum(1 for pred, gold in zip(predictions, gold_labels) if pred == gold) / len(predictions)
# For NLG benchmark (TranslatedDolly)
# Pairs of (model_output, baseline_output) are fed to Command R+ judge with a comparison prompt.
# Judge returns preference (model_wins, baseline_wins, or tie).
win_rates = sum(1 for decision in judge_decisions if decision == 'model_wins') / total_pairs * 100
```
## Common pitfalls
- Different shot settings are used across benchmarks (5-shot for MMLU, 6-shot for Taxi1500), which can affect few-shot performance comparability if not standardized.
- NLG evaluation relies on an LLM-as-a-judge (Command R+), which may introduce language-specific biases or inconsistencies in preference scoring across the 21 evaluated languages.
## Evidence (verbatim from paper)
> We evaluate the models in both multilingual and monolingual settings for NLU and open-ended generation tasks. Two evaluations use TranslatedDolly (Singh et al., [2024]), a translated version of Dolly (Conover et al., [2023]), a human-annotated English instruction-tuning dataset. Multilingual settings. NLU: Multilingual MMLU (Lai et al., [2023]) dataset, created by translating the English MMLU dataset to 31 languages. We evaluate using the lm-evaluation-harness framework EleutherAI ([2024]) with a 5-shot setup. NLG: TranslatedDolly, evaluated on 21 languages using the multilingual Command R+ Cohere ([2024]) model as an LLM judge. Monolingual low-resource settings. NLU: Taxi1500 (Ma et al., [2023]) for classification with a 6-shot setup based on a parallel Bible corpus covering 1500 languages. NLG: TranslatedDolly. Win rates are determined by Command R+ as judge.
## Citation
```bibtex
@misc{koksal2024muri,
title={MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions},
author={Köksal et al. (2024)},
year={2024},
note={arXiv:2409.12958}
}
```
- arXiv: 2409.12958
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!