Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill vcc2018-vc-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Vcc2018 Vc Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-vcc2018-vc-eval)More formats (shields.io, HTML) on the badges page.
---
name: vcc2018-vc-eval
description: Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity. Use when the user wants to benchmark on VCC 2018, or asks about evaluating this task. Reports MCD.
metadata:
skill_kind: dataset_eval
source_arxiv: 2102.12841
bibtex_key: kaneko2021maskcycleganvc
confidence: high
---
# vcc2018-vc-eval
> MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames — Kaneko et al. (2021) (arXiv:2102.12841, 2021)
## What this evaluates
Evaluates non-parallel voice conversion by measuring how well a model transforms a source speaker's speech into a target speaker's voice while preserving linguistic content. It assesses both acoustic fidelity using spectral distortion metrics and perceptual quality through human listening tests for naturalness and speaker similarity.
## Datasets
- **VCC 2018** — total ?; splits: train (81), test (35)
## Metrics
- `MCD` **(primary)** — range: other
- Mel-cepstral distortion calculated on 35-dimensional mel-cepstra extracted from the converted and target waveforms using the WORLD analyzer. Lower values indicate better spectral fidelity.
- `KDSD` — range: other
- Kernel DeepSpeech Distance measuring maximum mean discrepancy in the DeepSpeech2 feature space. Lower values indicate better correlation with human judgment.
## Input / output format
**Input**: 80-dimensional log mel-spectrogram extracted from source speaker audio (window length 1024, hop length 256 samples, sampled at 22.05 kHz).
**Output**: Converted 80-dimensional log mel-spectrogram, which is subsequently synthesized into a waveform using a pretrained MelGAN vocoder.
## Scoring recipe
```python
def compute_mcd(converted_waveform, target_waveform):
mc_pred = world_extract_mel_cepstrum(converted_waveform, dim=35)
mc_gold = world_extract_mel_cepstrum(target_waveform, dim=35)
return mean(sqrt(sum((mc_pred - mc_gold)**2, axis=-1)))
def compute_kdsd(converted_waveform, target_waveform):
feat_pred = deepspeech2_features(converted_waveform)
feat_gold = deepspeech2_features(target_waveform)
return maximum_mean_discrepancy(feat_pred, feat_gold)
# Both metrics are averaged across all test utterances and speaker pairs.
```
## Common pitfalls
- MCD and KDSD are inverse metrics where lower values indicate better performance, contrary to standard accuracy metrics.
- Direct frame-aligned comparison between converted and target spectrograms is invalid due to lack of alignment; feature-space metrics must be used instead.
- Subjective evaluation uses two different test formats: AB test for naturalness and XAB test for speaker similarity, which must not be conflated.
## Evidence (verbatim from paper)
> Because a direct comparison between the converted and target mel-spectrograms is difficult owing to the lack of a correct alignment, we used two metrics: (1) mel-cepstral distortion (MCD), which is the most commonly applied measure and calculates the distance within the mel-cepstral domain (particularly, a 35-dimensional mel-cepstrum was extracted from the converted or targeted waveform using the WORLD analyzer [50]), and (2) Kernel DeepSpeech Distance (KDSD) [51], which computes the maximum mean discrepancy within the DeepSpeech2 feature space [52] and is shown to be well correlated with human judgement [51]. For both metrics, the smaller the value, the better the performance.
## Citation
```bibtex
@misc{kaneko2021maskcycleganvc,
title={MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames},
author={Kaneko et al. (2021)},
year={2021},
note={arXiv:2102.12841}
}
```
- arXiv: 2102.12841
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!