Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill catwalk-default-metrics --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Catwalk Default Metrics?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-catwalk-default-metrics)More formats (shields.io, HTML) on the badges page.
---
name: catwalk-default-metrics
description: Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.
metadata:
skill_kind: metric
source_arxiv: 2312.10253
bibtex_key: groeneveld2023catwalk
confidence: high
---
# catwalk-default-metrics
> Catwalk: A Unified Language Model Evaluation Framework for Many Datasets — Groeneveld et al. (2023) (arXiv:2312.10253, 2023)
## What this evaluates
Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration.
## Datasets
- (no dataset; pure metric skill)
## Metrics
- `Accuracy & relative improvement over random baseline` **(primary)** — range: [0, 1] | percent
- Computes standard accuracy for multiple-choice, classification, and entailment tasks. Also calculates relative improvement over a uniform random baseline.
- `SQuAD metric` **(primary)** — range: [0, 1]
- Follows the exact implementation by Rajpurkar et al. (2016), computing token-level exact match and F1 score for open-ended question answering.
- `Perplexity per word, perplexity per byte, and entropy` — range: other
- Computes average log-probability per token (perplexity per word), per byte of text, and Shannon entropy over the predicted token distribution for language modeling tasks.
## Input / output format
**Input**: Dataset-specific inputs: multiple-choice prompts, classification/entailment pairs, open-ended QA questions, or raw text sequences for language modeling.
**Output**: Model predictions including class labels, generated text spans, or token-level probability distributions.
## Scoring recipe
```python
def compute_metrics(predictions, gold, dataset_type, model):
if dataset_type in ['multiple-choice', 'classification', 'entailment']:
acc = mean(predictions == gold)
random_acc = 1.0 / num_classes
return {'accuracy': acc, 'relative_improvement': (acc - random_acc) / random_acc}
elif dataset_type == 'qa':
return squad_metric(predictions, gold) # Rajpurkar et al. 2016
elif dataset_type == 'lm':
log_probs = model.log_prob(predictions)
return {
'perplexity_per_word': exp(-mean(log_probs)),
'perplexity_per_byte': exp(-mean(log_probs) / mean(len(token) for token in predictions)),
'entropy': -mean(sum(p * log(p) for p in dist))
}
```
## Common pitfalls
- Metrics are attached to datasets, not models; comparing models requires ensuring the same dataset-attached metrics are used.
- The Eleuther model style does not automatically compute suggested metrics, unlike other implemented styles.
- SQuAD metric refers to the specific Rajpurkar et al. (2016) implementation, which computes both exact match and F1, not just a single score.
## Evidence (verbatim from paper)
> By default, Catwalk computes the following metrics and can be easily extended to include others. - For multiple-choice, classification, and entailment datasets: Accuracy and relative improvement over the random baseline - For question-answering datasets: SQuAD metric as defined by Rajpurkar et al. (2016). - For language modeling datasets: Perplexity per word, perplexity per byte, and entropy.
## Citation
```bibtex
@misc{groeneveld2023catwalk,
title={Catwalk: A Unified Language Model Evaluation Framework for Many Datasets},
author={Groeneveld et al. (2023)},
year={2023},
note={arXiv:2312.10253}
}
```
- arXiv: 2312.10253
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!