Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill indicgenbench-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Indicgenbench Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-indicgenbench-eval)More formats (shields.io, HTML) on the badges page.
---
name: indicgenbench-eval
description: Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning. Use when the user wants to benchmark on CrossSum-In, Flores-In, XQuAD-In, XorQA-In, or asks about evaluating this task. Reports Character-F1 (ChrF), SQuAD-style Token-F1.
metadata:
skill_kind: dataset_eval
source_arxiv: 2404.16816
bibtex_key: singh2024indicgenbench
confidence: high
---
# indicgenbench-eval
> IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages — Singh et al. (2024) (arXiv:2404.16816, 2024)
## What this evaluates
Evaluates the multilingual and cross-lingual generation capabilities of LLMs across 29 Indic languages, covering summarization, machine translation, and question answering. It probes how model performance scales with language resourcedness, in-context learning, and fine-tuning.
## Datasets
- **CrossSum-In** — total ?; splits: test (-1)
- **Flores-In** — total ?; splits: test (-1)
- **XQuAD-In** — total ?; splits: test (-1)
- **XorQA-In** — total ?; splits: test (-1)
## Metrics
- `Character-F1 (ChrF)` **(primary)** — range: [0, 1]
- Computes n-gram overlap between reference and hypothesis at the character level, combining precision and recall into an F1 score. Used for summarization and translation tasks.
- `SQuAD-style Token-F1` **(primary)** — range: [0, 1]
- Computes token-level precision and recall between the predicted answer span and the gold answer, then returns their harmonic mean. Used for QA tasks.
## Input / output format
**Input**: Task-specific prompts in a few-shot or zero-shot setting. For summarization/translation: source text in one language, target in another. For QA: question and context passage in the target language or English.
**Output**: Generated text: a summary, a translated sentence, or a short answer span.
## Scoring recipe
```python
def compute_chrF(hypothesis, reference):
# Character n-gram overlap (typically n=6)
return chrF_score(hypothesis, reference)
def compute_token_f1(pred_span, gold_span):
pred_tokens = set(pred_span.split())
gold_tokens = set(gold_span.split())
if not pred_tokens or not gold_tokens:
return 0.0
precision = len(pred_tokens & gold_tokens) / len(pred_tokens)
recall = len(pred_tokens & gold_tokens) / len(gold_tokens)
return 2 * precision * recall / (precision + recall)
```
## Common pitfalls
- Token-level metrics like ROUGE and BLEU are unreliable for low-resource Indic languages; Character-F1 must be used instead.
- Translation performance differs significantly by direction (en→xx vs xx→en), with generation often lagging behind understanding.
- High token fertility in low-resource languages reduces the number of in-context examples that fit in the context window, artificially lowering few-shot performance.
## Evidence (verbatim from paper)
> For the cross-lingual summarization and translation tasks, CrossSum-In and Flores-In, we report Character-F1 (ChrF) metric since token-level metrics like ROUGE and BLEU are not reliable for low-resource languages. To stay consistent with existing literature on QA tasks, we report SQuAD-style Token-F1 on our XQuAD-In and XorQA-In QA tasks.
## Citation
```bibtex
@misc{singh2024indicgenbench,
title={IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages},
author={Singh et al. (2024)},
year={2024},
note={arXiv:2404.16816}
}
```
- arXiv: 2404.16816
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!