The TTS/STT/audio-playback system behind the one `speak()` entry point. Use when working on read-aloud, the Listen panel, voice/speed settings, transcription, mic capture, SpeakerButton, iOS playing silence, audio API routes, or files in features/audio/, features/tts/, hooks/tts/, app/api/audio/.
Installs into .claude/skills of the current project.
Are you the author of Tts Audio System?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/armanisadeghi-tts-audio-system)
---
name: tts-audio-system
description: "The TTS/STT/audio-playback system behind the one `speak()` entry point. Use when working on read-aloud, the Listen panel, voice/speed settings, transcription, mic capture, SpeakerButton, iOS playing silence, audio API routes, or files in features/audio/, features/tts/, hooks/tts/, app/api/audio/."
---
# TTS / Audio System
## π¨ Start here β the four things that are easy to get wrong
1. **`speak()` is the ONE entry point** (`features/audio/service/speak.ts`). Never
hand-roll a TTS call, never construct a player, never read a voice preference
yourself. React surfaces use `useSpeech()` on top of it. To let someone HEAR
a catalog model's voice (e.g. a speech-script speaker), call
`speak({ text, sample: { model, voice } })` β the catalog lane plays the
vendor's own sample (ElevenLabs) or a line the server renders once per
(model, voice) via `POST /audio/voice-preview`.
2. **Voice / speed / language come from the tiered `listening` config**
(`features/audio/service/listeningConfig.ts`) β system β org β user, user wins.
**Never read `userPreferences.voice.*` for playback**; those fields survive
only as the pre-fetch boot fallback. See the tiered-config section below.
3. **Audio output must be unlocked inside a user gesture** or iOS plays silence.
Call `primeAudioOutput()` (`features/audio/unlock.ts`) synchronously in any
handler that will later start audio. See the iOS section below.
4. **One voice at a time, app-wide.** `playbackLock` arbitrates; the
`playbackQueue` + `audioSessionRegistry` own what is playing and its history.
Producing audio without a session trips the runtime bypass guard.
## Architecture Overview
Two independent TTS engines, one STT engine, and a voice assistant pipeline.
### TTS Engines
| Engine | Provider | Transport | Primary Hook | API Route |
|--------|----------|-----------|-------------|-----------|
| **Cartesia** | Cartesia API | WebSocket (PCM F32LE, 44.1kHz) | `useCartesiaSpeaker` | `GET /api/cartesia` |
| **Groq/PlayAI** | Groq SDK | REST β WAV blob | `useTextToSpeech` | `POST /api/audio/text-to-speech` |
### STT Engine
| Route | Provider | Auth | Max Size |
|-------|----------|------|----------|
| `POST /api/audio/transcribe` | Groq Whisper | Supabase cookie/Bearer | 4.5 MB |
| `POST /api/audio/transcribe-url` | Groq Whisper | Supabase cookie/Bearer | 100 MB (URL) |
| `POST /api/audio/log-error` | Supabase | Supabase cookie/Bearer | N/A |
### Auth
All audio API routes use `resolveUser` from `utils/supabase/resolveUser.ts` β dual-mode: Supabase session cookie OR Bearer token. Both TTS routes (`/api/cartesia`, `/api/audio/text-to-speech`) require authentication.
---
## Directory Map
```
features/tts/ β Production TTS feature (Cartesia + Groq)
hooks/
useCartesiaSpeaker.ts β PRIMARY Cartesia hook (lazy, self-contained)
useTextToSpeech.ts β Groq/PlayAI hook (REST, WAV blob)
components/
SpeakerButton.tsx β Single play/pause toggle (lazy-loads core)
SpeakerButtonCore.tsx β Dynamically imported Cartesia logic
SpeakerGroup.tsx β 3-button group shell (Play, Pause, Stop)
SpeakerGroupCore.tsx β Dynamically imported 3-button core
SpeakerCompactGroup.tsx β 2-button group shell (Play/Pause, Stop)
SpeakerCompactGroupCore.tsx
AudioPlayerButton.tsx β Groq/PlayAI button (uses useTextToSpeech)
types.ts β Shared types (EnglishVoice, SpeakerVariant, etc.)
hooks/tts/ β Legacy/specialized hooks
simple/
useCartesiaControls.ts β Full controls (connect-on-mount, script state)
useCartesiaWithPreferences.ts β Redux voice prefs + connect-on-mount
useSimpleCartesia.ts β Minimal Cartesia (connect-on-mount)
useCartesia.ts β Legacy full-featured (direct API key, eager)
usePlayer.ts β Raw PCM stream player (Web Audio API, 24kHz)
usePlayerSafe.ts β Enhanced usePlayer with explicit init
useVoiceChat.ts β Voice assistant pipeline (VAD + STT + AI + TTS)
useVoiceChatCdn.ts β CDN variant of voice chat
useVoiceChatWithAutoSleep.ts β Auto-sleep extension
features/audio/ β THE modern audio system (start here, not features/tts/)
service/
speak.ts β π¨ THE app-wide "turn text into audio" entry point
useSpeech.ts β React face of speak() (per-surface status)
listeningConfig.ts β π¨ tiered voice/speed/language (systemβorgβuser)
useListeningSettings.ts β settings-pane face; update() writes MY tier
engines.ts β AV engine registry (voice lists live here ONCE)
unlock.ts β π¨ iOS/WebKit gesture unlock (see iOS section)
activation.ts β one-way latch that mounts the lazy audio system
playback/
playbackQueue.ts β the single app-wide queue (one at a time)
playbackLock.ts β app-wide single-output arbiter (start-always-wins)
AudioPlaybackHost.tsx β queue β Redux mirror; installs unlock listeners
adapters/cartesiaAdapter.ts β streaming PCM lane (SinkAwarePlayer)
adapters/catalogAdapter.ts β server-catalog lane (HTMLAudioElement)
session/
audioSessionRegistry.ts β ALL audio activity (in + out, live + history)
usePlaybackSessionController.ts / useMediaElementPlaybackSession.ts
bypassGuard.ts β screams when audio is produced with no session
sinkAwarePlayer.ts β our WebPlayer fork; output-device routing
hooks/ components/ services/ voice/ β recording, mic UI, voice selection UI
features/window-panels/windows/listen/
ListenSummaryWindow.tsx β the Listen panel (summarize-for-listening player)
providers/AudioSystemHost.tsx β the ONE ssr:false boundary (lazy mount)
providers/AudioOutputHostImpl.tsx β app-root streaming speaker owner
lib/cartesia/ β Low-level Cartesia client & types
client.ts β LEGACY: direct NEXT_PUBLIC_CARTESIA_API_KEY
cartesia.types.ts β Full Cartesia type definitions
voices.ts / voices.json β Voice catalog
app/api/cartesia/route.ts β Token endpoint (authenticated)
app/api/audio/ β TTS + STT API routes
utils/supabase/resolveUser.ts β Shared auth resolution
utils/markdown-processors/parse-markdown-for-speech.ts β MD β plain text
```
---
## Hook Selection Guide
**Adding TTS to a new component?** Use `useCartesiaSpeaker` or a Speaker* component.
| Need | Solution |
|------|----------|
| Single play/pause button | `<SpeakerButton text={content} />` |
| 3-button group (Play/Pause/Stop) | `<SpeakerGroup text={content} />` |
| 2-button group (toggle + Stop) | `<SpeakerCompactGroup text={content} />` |
| Programmatic TTS with Redux prefs | `useCartesiaSpeaker({ processMarkdown: true })` |
| Groq/PlayAI TTS (non-Cartesia) | `useTextToSpeech()` or `<AudioPlayerButton text={content} />` |
| Full voice config UI (demo/playground) | `useCartesiaControls()` or `useSimpleCartesia()` |
| Voice assistant with VAD | `useVoiceChat()` or `useVoiceChatWithAutoSleep()` |
### Why multiple hooks exist
The legacy hooks (`useCartesiaControls`, `useSimpleCartesia`, `useCartesia`) connect eagerly on mount and manage their own voice/emotion/speed state. They exist for playground and demo pages where the user needs full control over Cartesia parameters.
`useCartesiaSpeaker` is the production hook: lazy (does nothing until `speak()` is called), reads voice preferences from Redux, and has proper error handling. Always prefer it for new features.
---
## Speaker Component Contracts
All Speaker* components follow the same pattern:
1. **Thin shell** renders static disabled buttons (zero JS loaded)
2. **First click** triggers `React.lazy()` β dynamically imports Core
3. **Core** initializes `useCartesiaSpeaker` and begins playback
4. **Shape never changes** β buttons are always rendered, unavailable actions are disabled
### Props (all variants)
```typescript
interface SpeakerProps {
text: string; // Content to speak
processMarkdown?: boolean; // Strip markdown (default: true)
className?: string; // Applied to outer container
disabled?: boolean; // Disable all buttons
}
// SpeakerButton also accepts:
variant?: 'glass' | 'transparent' | 'solid' | 'group';
```
### Icon System
Speaker buttons use raw SVG icons from `@ai-matrx/tap-target/buttons`, NOT Lucide. Available: `PlayTapButton`, `PauseTapButton`, `StopTapButton`, `Volume2TapButton`. (Formerly `components/icons/tap-buttons.tsx`, deleted in the 2026-08-30 C9 package adoption.)
### Styling
Uses `TapTargetButton` / `TapTargetButtonGroup` from the `@ai-matrx/tap-target` package (C9 adoption, 2026-08-30 β the old in-repo copy is gone). Glass styling via `matrx-glass` CSS classes (globally available in `app/globals.css`).
---
## API Contracts
### GET /api/cartesia
**Auth:** Required (cookie or Bearer)
**Response:** `{ token: string }` β short-lived Cartesia access token
**Errors:** 401 (no auth), 500 (token generation failed)
### POST /api/audio/text-to-speech
**Auth:** Required (cookie or Bearer)
**Body:** `{ text: string; voice?: EnglishVoice; model?: string }`
**Response:** `audio/wav` binary (Cache-Control: 1 year)
**Limits:** text max 10,000 chars, voice must be valid PlayAI voice
**Errors:** 400 (validation), 401 (no auth), 429 (rate limit), 500
### POST /api/audio/transcribe
**Auth:** Required
**Body:** `multipart/form-data` with `file` (audio), optional `language`, `prompt`
**Response:** `{ success, text, language, duration, segments, _meta: { attempts } }`
**Limits:** 4.5 MB file size, allowed types: flac/mp3/mp4/mpeg/mpga/m4a/ogg/wav/webm
**Retries:** 3 with exponential backoff, logged to `audio_transcription_errors`
### POST /api/audio/transcribe-url
**Auth:** Required
**Body:** `{ url: string; language?; prompt? }` β URL must be on the route's allowlist (Python backend tiers + AWS S3 hostnames; signed cld_files URLs qualify)
**Response:** Same as /transcribe
**Limits:** 100 MB via Groq URL parameter
---
## Voice / speed / language β the tiered listening config (2026-08-28)
π¨ **Cartesia voice, speed, and language NO LONGER live in `userPreferences.voice` for playback.** They resolve through the tiered surface-config `listening` namespace (system β org β user, user wins) via `features/audio/service/listeningConfig.ts`:
- React reads: `selectListeningVoice/Speed/Language` (+ `selectListeningVoiceId(state, purpose)`).
- Imperative reads (`speak()`, playback adapters): `getListeningSettings()` β resolves at START time so replays honor current settings.
- Writes: `useListeningSettings().update(patch)` β ONE row at the user's tier (never a whole-preferences write).
- System default: platform-global row, admin-edited at `/administration/ui/surfaces/matrx-user/assistant-message` (Config namespaces). Org rows override; user rows win.
`userPreferences.voice.{voice,speed,language}` remain ONLY as the pre-fetch boot fallback; `emotion`/`wakeWord`/`microphone`/`speaker` still live there.
## Redux Voice Preferences (legacy fields β see above for voice/speed/language)
```typescript
// state.userPreferences.voice (Cartesia)
interface VoicePreferences {
voice: string; // LEGACY for playback β boot fallback only
language: string; // LEGACY for playback β boot fallback only
speed: number; // LEGACY for playback β boot fallback only
emotion: string;
microphone: boolean;
speaker: boolean;
wakeWord: string;
}
// state.userPreferences.textToSpeech (Groq/PlayAI)
interface TextToSpeechPreferences {
preferredVoice: GroqTtsVoice; // e.g. "Cheyenne-PlayAI"
autoPlay: boolean;
processMarkdown: boolean;
}
```
---
## π¨ iOS / WebKit β the output-unlock law (2026-08-30)
**Every browser on iPhone is WebKit β Chrome and Firefox included β so these
rules are not "a Safari edge case", they are the whole mobile platform.**
Two WebKit behaviors made all mobile audio silent, with no error:
1. An `AudioContext` created (or resumed) **outside a user gesture** starts
`suspended` and stays silent. Our TTS starts audio from websocket callbacks
seconds after the tap, so per-utterance contexts were always suspended.
2. Web Audio is muted by the **ringer/silent switch** unless the page declares
`navigator.audioSession.type = "playback"` (iOS 16.4+).
`features/audio/unlock.ts` is the fix and the only place this logic lives:
- `installAudioUnlockListeners()` β capture-phase `pointerdown`/`keydown`/
`touchend` listeners; mounted ONCE by `AudioPlaybackHost`. Unlocks on the
user's first interaction anywhere, so the page is usually already unlocked
before any audio is requested.
- `primeAudioOutput()` β **call this synchronously in every handler that will
later start audio.** Idempotent and cheap. Already called by
`enqueuePlayback` / `resumePlayback` / `playPlaybackItem`, both Listen
actions, and the Listen panel transport.
- `getUnlockedAudioContext()` / `getPrimedMediaElement()` β the shared,
gesture-unlocked handles. `SinkAwarePlayer` schedules into the shared
context when it exists (per-utterance `GainNode`; **never closes it**);
`catalogAdapter` plays through the shared element.
**Rules when you touch audio:**
- Adding a new "start audio" button? Call `primeAudioOutput()` in the handler.
- Never mint a bare `new AudioContext()` for playback β go through the queue.
- Never close the shared context (`stop()` in shared mode tears down the
utterance only).
- A context that will not reach `running` **throws a user-facing error**
("The browser blocked audio output β tap the play button to start sound").
Keep it that way: silence with no error is the bug this replaced.
---
## The Listen panel β summarize-for-listening
`features/window-panels/windows/listen/ListenSummaryWindow.tsx` (overlay
`listenSummaryWindow`). Two entry points, both in one "Listen" context-menu
submenu and in the assistant action bar's β― menu:
- **Summarize without playing** β summary streams in as text; user presses Play.
- **Summarize & listen** β stream-to-stream: the summary is spoken as it is
written, via the app-root speaker (`voicePlaybackBus` request with
`includeActive: true`).
The summarizing agent is resolved from the `spoken_summary` surface role,
falling back to the platform home surface (`LISTENING_HOME_SURFACE`) whose
`ambient.spoken_summary` mandate holds the default agent β so it works on
every surface for every user with no personal binding. **Never put an agent
UUID in code here.** Menu wiring: `features/context-menu-v3/` (`listen`
submenu role) + `messageActionRegistry.listeningItems`.
---
## Known Deferred Issues
These are documented design decisions, not bugs:
1. **Five Cartesia hook variants** β Consolidation into one hook requires updating ~15 consumer files. Deferred as separate effort. Use `useCartesiaSpeaker` for all new code.
2. **`lib/cartesia/client.ts` leaks API key** β Uses `NEXT_PUBLIC_CARTESIA_API_KEY`. Legacy hook `useCartesia` depends on it. Will be removed when legacy hook is consolidated.
3. **No connection health monitoring** β No heartbeat/ping. Silent WebSocket drops discovered only on next `speak()`.
4. **No playback progress for Cartesia** β the catalog (element) lane tracks duration/currentTime; the streaming PCM lane does not.
5. **Token expiry** β Cartesia token fetched once, no refresh. If it expires mid-session, `connectCartesiaTts` refreshes and retries once automatically.
**Two entries were REMOVED on 2026-09-08 because they became false** β do not
reinstate them from an old copy of this file: *"no global TTS instance
management"* (the `playbackQueue` + `playbackLock` + `audioSessionRegistry`
trio has been the single owner since the unified-audio work) and *"no abort
for in-flight TTS"* (`skipPlayback` / `clearPlayback` / a session's `stop`
control all abort; `playbackLock` preempts across paths).
---
## Adding a New Speaker Variant
1. Create `SpeakerNewVariant.tsx` (thin shell) in `features/tts/components/`
2. Create `SpeakerNewVariantCore.tsx` (default export, uses `useCartesiaSpeaker`)
3. Use `React.lazy()` in shell to import core
4. Use `TapTargetButton*` from `@ai-matrx/tap-target` (icons: `@ai-matrx/tap-target/buttons`)
5. Export from `features/tts/components/index.ts`
6. Never hide buttons β disable unavailable actions
7. Never change component shape during state transitions
## Modifying API Routes
All audio API routes use `resolveUser` from `utils/supabase/resolveUser.ts`. When adding new routes:
1. Import and call `resolveUser(request)` first
2. Return 401 if `!user`
3. Use structured error responses with `code` field
4. For Groq routes: implement retry with exponential backoff
5. Log errors to `audio_transcription_errors` via `logTranscriptionError()`