Evaluates retrieval over long multilingual documents (up to 8,192 tokens), testing a model's ability to capture information from extended contexts across multiple languages. Use when the user wants to benchmark on MLDR, or asks about evaluating this task. Reports nDCG@10.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill mldr-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mldr Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-mldr-eval)More formats (shields.io, HTML) on the badges page.
---
name: mldr-eval
description: Evaluates retrieval over long multilingual documents (up to 8,192 tokens), testing a model's ability to capture information from extended contexts across multiple languages. Use when the user wants to benchmark on MLDR, or asks about evaluating this task. Reports nDCG@10.
metadata:
skill_kind: dataset_eval
source_arxiv: 2402.03216
bibtex_key: chen2024m3embedding
confidence: high
---
# mldr-eval
> M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation — Chen et al. (2024) (arXiv:2402.03216, 2024)
## What this evaluates
Evaluates retrieval over long multilingual documents (up to 8,192 tokens), testing a model's ability to capture information from extended contexts across multiple languages.
## Datasets
- **MLDR** — total ?; splits: test (-1)
## Metrics
- `nDCG@10` **(primary)** — range: [0, 1]
- Normalized Discounted Cumulative Gain at rank 10, measuring the quality of ranked long-document retrieval results.
## Input / output format
**Input**: Long multilingual documents (from Wikipedia, Wudao, mC4) and corresponding queries.
**Output**: A ranked list of retrieved long documents.
## Scoring recipe
```python
def compute_ndcg_at_10(retrieved_ids, relevant_ids):
dcg = 0.0
for i, doc_id in enumerate(retrieved_ids[:10]):
if doc_id in relevant_ids:
dcg += 1.0 / math.log2(i + 2)
idcg = sum(1.0 / math.log2(i + 2) for i in range(min(len(relevant_ids), 10)))
return dcg / idcg if idcg > 0 else 0.0
```
## Common pitfalls
- Sparse retrieval surprisingly outperforms dense retrieval on this benchmark, contrary to typical dense-retrieval assumptions.
- The model's max length is fixed at 8192 tokens, which may truncate longer documents if not handled.
## Evidence (verbatim from paper)
> We evaluate the retrieval performance with longer sequences with two benchmarks: MLDR (Multilingual Long-Doc Retrieval), which is curated by the multilingual articles from Wikipedia, Wudao and mC4... measured by nDCG@10.
## Citation
```bibtex
@misc{chen2024m3embedding,
title={M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
author={Chen et al. (2024)},
year={2024},
note={arXiv:2402.03216}
}
```
- arXiv: 2402.03216
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!