Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard, MT-Bench, MATH, GSM-8K, HumanEval, MBPP, MMLU, LUC-EVAL, or asks about evaluating this task. Reports Overall*.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill infinity-instruct-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Infinity Instruct Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-infinity-instruct-eval)More formats (shields.io, HTML) on the badges page.
---
name: infinity-instruct-eval
description: Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard, MT-Bench, MATH, GSM-8K, HumanEval, MBPP, MMLU, LUC-EVAL, or asks about evaluating this task. Reports Overall*.
metadata:
skill_kind: dataset_eval
source_arxiv: 2506.11116
bibtex_key: li2025infinityinstruct
confidence: high
---
# infinity-instruct-eval
> Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models — Li et al. (2025) (arXiv:2506.11116, 2025)
## What this evaluates
Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks.
## Datasets
- **AlpacaEval 2.0** — total ?; splits: test (-1)
- **Arena-Hard** — total ?; splits: test (-1)
- **MT-Bench** — total ?; splits: test (-1)
- **MATH** — total ?; splits: test (-1)
- **GSM-8K** — total ?; splits: test (-1)
- **HumanEval** — total ?; splits: test (-1)
- **MBPP** — total ?; splits: test (-1)
- **MMLU** — total ?; splits: test (-1)
- **LUC-EVAL** — total ?; splits: test (-1)
## Metrics
- `AlpacaEval 2.0` — range: percent
- Win rate percentage against a reference model (typically GPT-4-turbo) using LLM-as-judge.
- `Arena-Hard` — range: percent
- Average score from pairwise comparisons judged by LLMs, normalized to 0-100.
- `MT-Bench` — range: percent
- Average score across multi-turn conversations rated by LLMs, mapped to 0-100 for reporting.
- `MATH` — range: percent
- Exact-match accuracy on mathematical reasoning problems.
- `GSM-8K` — range: percent
- Exact-match accuracy on grade-school math word problems.
- `HumanEval` — range: percent
- Pass@1 accuracy based on unit test execution.
- `MBPP` — range: percent
- Pass@1 accuracy based on unit test execution for code generation.
- `MMLU` — range: percent
- Exact-match accuracy on multiple-choice questions across 57 subjects.
- `LUC-EVAL` — range: percent
- Accuracy on a specialized evaluation suite for foundational capabilities.
- `Overall*` **(primary)** — range: percent
- Average of normalized scores across all listed benchmarks. MT-Bench scores are explicitly mapped to 0-100 before averaging.
## Input / output format
**Input**: Instruction prompts or conversational turns from evaluation benchmarks.
**Output**: Model-generated text responses.
## Scoring recipe
```python
def compute_metrics(predictions, golds, benchmark):
if benchmark in ['AlpacaEval 2.0', 'Arena-Hard', 'MT-Bench']:
score = llm_judge_score(predictions, golds)
else:
score = exact_match_or_execution_accuracy(predictions, golds)
return score
def compute_overall(predictions, golds, benchmarks):
normalized_scores = []
for b in benchmarks:
s = compute_metrics(predictions, golds, b)
if b == 'MT-Bench':
s = s * 10 # Map 0-10 to 0-100
normalized_scores.append(s)
return sum(normalized_scores) / len(normalized_scores)
```
## Common pitfalls
- MT-Bench scores are originally on a 0-10 scale but must be multiplied by 10 to match the 0-100 scale of other benchmarks in the overall average.
- AlpacaEval 2.0 reports win rates against a specific reference model (usually GPT-4-turbo), so scores are not absolute but relative.
- Arena-Hard and MT-Bench rely on LLM-as-judge evaluations, which can introduce bias depending on the judge model and prompt.
## Evidence (verbatim from paper)
> To calculate the overall average score, we mapped MT-Bench scores to 0-100.
## Citation
```bibtex
@misc{li2025infinityinstruct,
title={Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models},
author={Li et al. (2025)},
year={2025},
note={arXiv:2506.11116}
}
```
- arXiv: 2506.11116
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!