Improve long-form text-to-speech through a running Voicebox REST server by planning semantic speech segments, generating them separately, preserving provenance, and merging compatible WAV audio. Use when a user explicitly asks to use Voicebox, a local cloned voice, or an authorized Voicebox profile for a long narration, article, script, voiceover, or multi-paragraph TTS asset. Do not use for generic TTS provider selection or ChatCut built-in cloud voices unless the user specifically chooses V...
Installs into .claude/skills of the current project.
Are you the author of Voicebox Longform Tts?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/cloudyview-voicebox-longform-tts)
---
name: voicebox-longform-tts
description: Improve long-form text-to-speech through a running Voicebox REST server by planning semantic speech segments, generating them separately, preserving provenance, and merging compatible WAV audio. Use when a user explicitly asks to use Voicebox, a local cloned voice, or an authorized Voicebox profile for a long narration, article, script, voiceover, or multi-paragraph TTS asset. Do not use for generic TTS provider selection or ChatCut built-in cloud voices unless the user specifically chooses Voicebox.
---
# Voicebox Long-form TTS
Render long narration as short, coherent speech beats instead of one unstable model call. Use the bundled standard-library script for API calls, polling, audio assembly, and the run manifest.
## Core rules
- Obtain explicit approval before generating a cloned or identity-sensitive voice. Pass `--authorized` only when that approval exists in the current task.
- Keep Voicebox loopback-only by default. Do not pass `--allow-remote` without approval to transmit the narration to that endpoint.
- Call the Voicebox REST API. Never read or modify its SQLite database.
- Treat a saved profile as voice identity, not guaranteed emotion or delivery.
- Split by meaning and performance intent. Prefer one primary speaking intention per segment.
- Remove direction markup such as `【语气:笃定】` from spoken text. Store direction in the plan's `intent` field.
- Do not assume inline emotion tags or `instruct` work. Forward `instruct` only after verifying the selected engine/backend supports it.
- Preserve every segment, generation ID, seed, hash, and final manifest. Never overwrite an existing run unless the user explicitly requests it.
- Treat technical QA as necessary but not sufficient. Ask the user to judge identity fidelity, delivery, emphasis, and joins by listening.
## Workflow
1. Resolve the bundled script relative to this `SKILL.md`:
```bash
python3 scripts/voicebox_longform.py health
python3 scripts/voicebox_longform.py profiles
```
2. Confirm the target profile and voice authorization. For cloned profiles, stop if authorization is absent.
3. Create an editable plan before rendering important narration:
```bash
python3 scripts/voicebox_longform.py plan \
--text-file narration.txt \
--output narration.plan.json
```
4. Review the plan. Keep most Chinese segments around 20–50 characters, allow up to about 70 when a thought must stay intact, and fill `intent` with concise non-spoken direction. Rewrite segment boundaries when the automatic split cuts meaning badly.
5. Render through an explicit profile:
```bash
python3 scripts/voicebox_longform.py render \
--plan narration.plan.json \
--profile "My authorized voice" \
--authorized \
--output narration.wav
```
6. Inspect the printed result and adjacent `narration.manifest.json`. Verify status, segment count, duration, sample rate, channels, peak level, and SHA-256. Preserve the `.voicebox-run/` directory.
7. Present the final WAV for listening. If one performance beat is weak, adjust that beat's text/intent and seed in a new version rather than editing the previous output in place.
## Planning expressive text
Convert human-friendly annotations into plan metadata:
```text
【语气:好奇;句尾上扬;重音:完全不会剪辑】
如果我完全不会剪辑,能不能直接告诉 AI 我想要什么?
```
Plan representation:
```json
{
"index": 1,
"text": "如果我完全不会剪辑,能不能直接告诉 AI 我想要什么?",
"intent": "好奇;句尾上扬;重音:完全不会剪辑"
}
```
The current script records `intent`; it does not speak it. See [references/api-contract.md](references/api-contract.md) for the plan schema, environment variables, Voicebox endpoints, and backend limitations.
## Output contract
For `narration.wav`, expect:
```text
narration.wav
narration.manifest.json
narration.voicebox-run/
├── run.json
├── segment-001.wav
└── ...
```
The tool refuses to overwrite these paths unless `--force` is explicitly supplied.