This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill triviaqa-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Triviaqa Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-triviaqa-eval)More formats (shields.io, HTML) on the badges page.
---
name: triviaqa-eval
description: This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains. Use when the user wants to benchmark on TriviaQA, or asks about evaluating this task. Reports exact match (EM).
metadata:
skill_kind: dataset_eval
source_arxiv: 1705.03551
bibtex_key: joshi2017triviaqa
confidence: high
---
# triviaqa-eval
> TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension — Mandar Joshi et al. (arXiv:1705.03551, 2017)
## What this evaluates
This benchmark evaluates reading comprehension on complex, compositional trivia questions that require multi-sentence reasoning and handling high lexical variability. It tests a model's ability to locate and extract precise answers from large, noisy evidence documents across different domains.
## Datasets
- **TriviaQA** — total ?; splits: train (-1), dev (-1), test (-1)
## Metrics
- `exact match (EM)` **(primary)** — range: [0, 1]
- Exact match (EM) and F1 over words in the answer(s). For questions that have Numerical and FreeForm answers, a single given answer is used as ground truth. For questions that have Wikipedia entities as answers, Wikipedia aliases are valid alongside the given answer.
- `F1` — range: [0, 1]
- Token-level F1 score computed between the predicted answer and the set of valid ground truth answers (including aliases).
## Input / output format
**Input**: A natural language question paired with one or more evidence documents (from either the Wikipedia or Web domain).
**Output**: An extracted answer span or entity string.
## Scoring recipe
```python
def compute_metrics(predictions, gold_answers):
em_scores = []
f1_scores = []
for pred, golds in zip(predictions, gold_answers):
valid_answers = set(golds) | set(golds_aliases)
em = 1.0 if pred in valid_answers else 0.0
em_scores.append(em)
pred_tokens = set(pred.split())
gold_tokens = set()
for g in valid_answers:
gold_tokens |= set(g.split())
if not pred_tokens or not gold_tokens:
f1 = 0.0
else:
intersection = len(pred_tokens & gold_tokens)
precision = intersection / len(pred_tokens)
recall = intersection / len(gold_tokens)
f1 = 2 * precision * recall / (precision + recall)
f1_scores.append(f1)
return sum(em_scores) / len(em_scores), sum(f1_scores) / len(f1_scores)
```
## Common pitfalls
- Evaluation granularity differs by domain: question-level accuracy is reported for Wikipedia, while document-level accuracy is reported for the Web domain due to high information redundancy.
- Oracle score upper bounds are artificially capped below 100% because training documents are truncated to the first 400 tokens to manage training time.
- Entity answers require matching against a set of Wikipedia aliases, not just the exact ground truth string.
## Evidence (verbatim from paper)
> We use the same evaluation metrics as SQuAD – exact match (EM) and F1 over words in the answer(s). For questions that have Numerical and FreeForm answers, we use a single given answer as ground truth. For questions that have Wikipedia entities as answers, we use Wikipedia aliases as valid answer along with the given answer.
## Citation
```bibtex
@misc{joshi2017triviaqa,
title={TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension},
author={Mandar Joshi et al.},
year={2017},
note={arXiv:1705.03551}
}
```
- arXiv: 1705.03551
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!