Evaluates zero-shot and one-shot text-to-speech voice cloning fidelity, multilingual synthesis capability, and cross-lingual generalization. It measures perceptual naturalness and speaker identity preservation through objective transcription and embedding similarity metrics, alongside human preference rankings. Use when the user wants to benchmark on Seed-TTS-eval, Artificial Arena, MiniMax Multilingual Test Set, or asks about evaluating this task. Reports WER, SIM.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill minimax-speech-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Minimax Speech Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-minimax-speech-eval)More formats (shields.io, HTML) on the badges page.
---
name: minimax-speech-eval
description: Evaluates zero-shot and one-shot text-to-speech voice cloning fidelity, multilingual synthesis capability, and cross-lingual generalization. It measures perceptual naturalness and speaker identity preservation through objective transcription and embedding similarity metrics, alongside human preference rankings. Use when the user wants to benchmark on Seed-TTS-eval, Artificial Arena, MiniMax Multilingual Test Set, or asks about evaluating this task. Reports WER, SIM.
metadata:
skill_kind: dataset_eval
source_arxiv: 2505.07916
bibtex_key: zhang2025minimaxspeech
confidence: high
---
# minimax-speech-eval
> MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder — Bowen Zhang et al. (2025) (arXiv:2505.07916, 2025)
## What this evaluates
Evaluates zero-shot and one-shot text-to-speech voice cloning fidelity, multilingual synthesis capability, and cross-lingual generalization. It measures perceptual naturalness and speaker identity preservation through objective transcription and embedding similarity metrics, alongside human preference rankings.
## Datasets
- **Seed-TTS-eval** — total 3000; splits: test-zh (2000), test-en (1000)
- **Artificial Arena** — total ?; splits: test (-1); repo https://artificialanalysis.ai
- **MiniMax Multilingual Test Set** — total 2400; splits: test (2400)
## Metrics
- `WER` **(primary)** — range: percent
- Word Error Rate computed by transcribing synthesized audio with Whisper-large-v3 (English) or Paraformer-zh (Chinese) and comparing against ground-truth text.
- `SIM` **(primary)** — range: [0, 1]
- Cosine similarity between speaker embeddings extracted from reference and synthesized audio using a WavLM-large model fine-tuned for speaker verification.
- `ELO` — range: other
- Dynamic ranking score derived from pairwise human preference judgments on the Artificial Arena leaderboard.
## Input / output format
**Input**: Target text string and a reference audio clip (untranscribed for zero-shot, optionally transcribed/prompted for one-shot).
**Output**: Synthesized audio waveform corresponding to the target text.
## Scoring recipe
```python
# WER
asr = WhisperLargeV3() if lang == 'en' else ParaformerZh()
pred_text = asr.transcribe(generated_audio)
wer = compute_wer(pred_text, ground_truth_text)
# SIM
encoder = WavLMLargeFineTuned()
ref_emb = encoder.encode(reference_audio)
gen_emb = encoder.encode(generated_audio)
sim = cosine_similarity(ref_emb, gen_emb)
# ELO
elo = artificial_arena_leaderboard.get_elo(model_name)
```
## Common pitfalls
- WER values are highly dependent on the specific ASR model used for transcription (Whisper-large-v3 vs Paraformer-zh), making cross-paper comparisons difficult.
- SIM scores use a custom fine-tuned WavLM-large speaker verification model, not the standard pre-trained WavLM, so results will not match baseline WavLM similarity scores.
- Artificial Arena ELO scores are dynamic and context-dependent, changing as new models are added or removed from the leaderboard.
## Evidence (verbatim from paper)
> The fidelity of voice cloning was quantitatively assessed using WER and SIM metrics on the Seed-TTS-eval (Anastassiou et al., 2024) test set. This dataset comprises two distinct subsets: test-zh (approximately 2,000 Chinese samples) and test-en (approximately 1,000 English samples). Each sample in these subsets includes a reference audio and a corresponding ground-truth audio from the identical speaker. For WER computation, synthesized English and Chinese audio were transcribed using Whisper-large-v3 (Radford et al., 2023) and Paraformer-zh (Gao et al., 2023b), respectively. SIM was determined by calculating the cosine similarity between speaker embeddings, which were extracted using a speaker verification model fine-tuned on WavLM-large.
## Citation
```bibtex
@misc{zhang2025minimaxspeech,
title={MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder},
author={Bowen Zhang et al. (2025)},
year={2025},
note={arXiv:2505.07916}
}
```
- arXiv: 2505.07916
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!