Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill jamendo-mt-qa-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Jamendo Mt Qa Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-jamendo-mt-qa-eval)More formats (shields.io, HTML) on the badges page.
---
name: jamendo-mt-qa-eval
description: Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.
metadata:
skill_kind: dataset_eval
source_arxiv: 2604.09721
bibtex_key: koh2026jamendomtqa
confidence: high
---
# jamendo-mt-qa-eval
> Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering — Koh et al. (2026) (arXiv:2604.09721, 2026)
## What this evaluates
Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination.
## Datasets
- **Jamendo-MT-QA** — total 36519; splits: full (36519), subset (2010)
## Metrics
- `accuracy` **(primary)** — range: [0, 1]
- 1 if the predicted answer exactly matches the ground truth, else 0. Applied to yes/no and short-answer questions.
- `BLEU` — range: [0, 1]
- N-gram precision with brevity penalty to measure surface-form similarity between predicted and gold sentences.
- `ROUGE-1/2/L` — range: [0, 1]
- Recall-oriented n-gram overlap metrics measuring surface-form similarity.
- `BERTScore` — range: [0, 1]
- Cosine similarity between contextual embeddings of predicted and gold tokens, averaged across tokens.
- `LLM-as-a-Judge score` **(primary)** — range: [1, 5]
- 1–5 rating assigned by an LLM based on correctness and comparative soundness using a fixed rubric.
## Input / output format
**Input**: Two audio tracks (or their generated captions) and a comparative question.
**Output**: A text answer: 'yes'/'no', a track identifier, or a natural language sentence.
## Scoring recipe
```python
def score(pred, gold, q_type):
if q_type == 'yes/no':
return 1.0 if pred.strip().lower() == gold.strip().lower() else 0.0
elif q_type == 'short':
return 1.0 if pred.strip() == gold.strip() else 0.0
elif q_type == 'sentence':
bleu = compute_bleu(pred, gold)
rouge = compute_rouge(pred, gold)
bert = compute_bertscore_f1(pred, gold)
llm_score = llm_judge(pred, gold, prompt, rubric) # 1-5
return {'bleu': bleu, 'rouge': rouge, 'bert_f1': bert, 'llm_judge': llm_score}
```
## Common pitfalls
- Relying solely on BLEU/ROUGE for sentence-level answers underestimates model capability due to low lexical overlap in open-ended generation.
- Comparing caption-based and multi-audio baselines directly without accounting for the intermediate abstraction step can misattribute performance gains.
- LLM-as-a-Judge scores are highly sensitive to the specific prompt and scoring rubric used; results are not directly comparable across different judge configurations.
## Evidence (verbatim from paper)
> For yes/no questions, we report accuracy. For short-answer questions, accuracy is computed by exact match between the predicted track identifier and the ground-truth answer. For sentence-level questions, we report BLEU and ROUGE-1,2,L as surface-form similarity metrics based on n-gram overlap, and BERTScore as a semantic similarity metric.
## Citation
```bibtex
@misc{koh2026jamendomtqa,
title={Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering},
author={Koh et al. (2026)},
year={2026},
note={arXiv:2604.09721}
}
```
- arXiv: 2604.09721
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!