Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity. Use when the user wants to benchmark on GTSinger, or asks about evaluating this task. Reports MOS-Q.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill gtsinger-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Gtsinger Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-gtsinger-eval)More formats (shields.io, HTML) on the badges page.
---
name: gtsinger-eval
description: Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity. Use when the user wants to benchmark on GTSinger, or asks about evaluating this task. Reports MOS-Q.
metadata:
skill_kind: dataset_eval
source_arxiv: 2409.13832
bibtex_key: zhang2024gtsinger
confidence: high
---
# gtsinger-eval
> GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks — Yu Zhang et al. (2024) (arXiv:2409.13832, 2024)
## What this evaluates
Evaluates singing voice synthesis models on technique-controllable generation, singer similarity, and audio quality across multiple languages and vocal techniques. It probes the model's ability to accurately control specific singing techniques (e.g., vibrato, mixed voice) while maintaining naturalness and timbre fidelity.
## Datasets
- **GTSinger** — total ?; splits: test (-1); repo https://github.com/AaronZ345/GTSinger
## Metrics
- `MOS-Q` **(primary)** — range: other
- Average Likert scale score (1-5) assigned by professional listeners evaluating synthesis quality (clarity, naturalness, and rich stylistic details).
- `MOS-S` — range: other
- Average Likert scale score (1-5) assigned by professional listeners evaluating singer similarity (timbre and style match to reference audio).
- `MOS-C` — range: other
- Average Likert scale score (1-5) assigned by professional listeners evaluating technique controllability (accuracy and expressiveness of technique control).
- `FFE` — range: other
- F0 Frame Error combining voicing decision error and F0 error to comprehensively capture pitch and voicing information.
- `MCD` — range: other
- Mean Cepstral Distortion: (10/ln10) * sqrt(2 * sum_{d=1}^D (c_t(d) - c_hat_t(d))^2), measuring spectral distortion between target and predicted MFCCs.
- `Cos` — range: [-1, 1]
- Cosine similarity between WavLM speaker verification embeddings extracted from synthesized and ground truth audio.
## Input / output format
**Input**: Synthesized singing audio and reference (ground truth) audio; for technique control, binary sequences indicating technique presence per phoneme.
**Output**: Subjective: Likert scores 1-5 per listener. Objective: FFE, MCD, and Cosine Similarity values.
## Scoring recipe
```python
def compute_metrics(predictions, gold, listeners):
# Subjective MOS-Q/S/C
scores = [listener.rating for listener in listeners] # 1-5
mos = sum(scores) / len(scores)
# Objective: MCD
mcd = (10 / ln(10)) * sqrt(2 * sum((c_t[d] - c_hat_t[d])**2 for d in range(D)))
# Objective: FFE
ffe = voicing_error_rate + f0_error_rate # Combined metric
# Objective: Cosine Similarity
cos_sim = cosine_similarity(wavlm_embed(predictions), wavlm_embed(gold))
return {'mos': mos, 'mcd': mcd, 'ffe': ffe, 'cos': cos_sim}
```
## Common pitfalls
- Listeners may conflate synthesis quality with singer similarity or technique controllability if instructions are not strictly enforced.
- FFE combines voicing and pitch errors, which can obscure whether performance issues stem from pitch accuracy or voicing detection.
- Subjective evaluation relies on a small test subset (50 sentences) and requires professional listeners, limiting reproducibility and scalability.
## Evidence (verbatim from paper)
> For each task, we randomly select 50 sentences from our test set for subjective evaluation... we conduct the MOS (Mean opinion score) evaluation... listeners are instructed to concentrate on synthesis quality... MOS-S evaluations, the listeners are directed to assess singer similarity... MOS-C, the listeners are informed to evaluate technique controllability... grade various singing voice samples on a Likert scale ranging from 1 to 5. ... we use F0 Frame Error (FFE)... employ Mean Cepstral Distortion (MCD) for measuring audio quality as the formula: MCD = (10/ln10) sqrt(2 sum (c_t(d) - c_hat_t(d))^2)... Cosine Similarity (Cos) is utilized to quantify the resemblance between the synthesized and reference singing voices.
## Citation
```bibtex
@misc{zhang2024gtsinger,
title={GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks},
author={Yu Zhang et al. (2024)},
year={2024},
note={arXiv:2409.13832}
}
```
- arXiv: 2409.13832
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!