This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill libribrain-speech-decoding-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Libribrain Speech Decoding Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-libribrain-speech-decoding-eval)More formats (shields.io, HTML) on the badges page.
---
name: libribrain-speech-decoding-eval
description: This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli. Use when the user wants to benchmark on LibriBrain, or asks about evaluating this task. Reports Balanced Accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2506.02098
bibtex_key: ozdogan2025libribrain
confidence: high
---
# libribrain-speech-decoding-eval
> LibriBrain: Over 50 Hours of Within-Subject MEG to Improve Speech Decoding Methods at Scale — Özdogan et al. (2025) (arXiv:2506.02098, 2025)
## What this evaluates
This benchmark evaluates non-invasive brain-computer interface (BCI) capabilities by testing neural speech decoding from magnetoencephalography (MEG) recordings. It probes a model's ability to detect speech presence, classify phonemes, and identify words from high-fidelity, within-subject neural data aligned with naturalistic audio stimuli.
## Datasets
- **LibriBrain** — total ?; splits: train (-1), val (-1), test (-1); repo https://github.com/neural-processing-lab/libribrain-experiments
## Metrics
- `Balanced Accuracy` **(primary)** — range: [0, 1]
- The macro average of recall across all classes. Computed as the mean of per-class true positive rates, ensuring equal weight to each class regardless of frequency.
- `F1-Score` — range: [0, 1]
- The harmonic mean of precision and recall. Reported as both micro- and macro-averaged values for multi-class tasks.
- `AUROC` — range: [0, 1]
- Area Under the Receiver Operating Characteristic curve, measuring the model's ability to distinguish between classes across all classification thresholds.
- `Jaccard Index` — range: [0, 1]
- Intersection over union (IoU) of predicted and ground truth positive sets.
- `Cross Entropy Loss` — range: other
- Negative log-likelihood of the true labels given the predicted probability distribution.
- `Top-10 Balanced Accuracy` — range: [0, 1]
- Macro average of balanced accuracy computed over the top-10 most frequent words in the vocabulary.
## Input / output format
**Input**: A segment or short window of MEG sensor recordings, temporally aligned to either a continuous audio segment (speech detection) or the onset of a specific phoneme/word in the stimulus audio.
**Output**: Predicted class label (e.g., speech/no-speech, one of 39 ARPAbet phonemes, or one of 250 target words) or a probability distribution over classes.
## Scoring recipe
```python
def compute_metrics(y_true, y_pred, num_classes):
recalls = []
precisions = []
for c in range(num_classes):
tp = np.sum((y_true == c) & (y_pred == c))
fn = np.sum((y_true == c) & (y_pred != c))
fp = np.sum((y_true != c) & (y_pred == c))
recalls.append(tp / (tp + fn) if (tp + fn) > 0 else 0.0)
precisions.append(tp / (tp + fp) if (tp + fp) > 0 else 0.0)
balanced_acc = np.mean(recalls)
f1_macro = np.mean([2 * p * r / (p + r) if (p + r) > 0 else 0.0
for p, r in zip(precisions, recalls)])
return balanced_acc, f1_macro
```
## Common pitfalls
- Macro F1-score is heavily penalized by the power-law distribution of phoneme frequencies, causing models to ignore rare phonemes and appear worse than they are on frequent classes.
- The random baseline for word classification is fixed at 1/250 (0.04), not uniform over the entire vocabulary or all possible words.
- Statistical significance is evaluated using exact permutation tests (1,024 sign-flips) rather than standard parametric tests, which must be replicated for valid comparison.
- Performance scales logarithmically with training data volume, not linearly, so small dataset size changes yield diminishing returns.
## Evidence (verbatim from paper)
> We assess model performance using a number metrics: F1-Score, Balanced Accuracy, Area Under the Receiver Operating Characteristic curve (AUROC), Jaccard Index, and Cross Entropy Loss. We present the results in Table 3.
## Citation
```bibtex
@misc{ozdogan2025libribrain,
title={LibriBrain: Over 50 Hours of Within-Subject MEG to Improve Speech Decoding Methods at Scale},
author={Özdogan et al. (2025)},
year={2025},
note={arXiv:2506.02098}
}
```
- arXiv: 2506.02098
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!