Local speech-to-text using Vosk. Lightweight, fast, fully offline. Perfect for transcribing Telegram voice messages, audio files, or any speech-to-text task without cloud APIs.
Pro shows the line behind each finding and how to fix it
Scanned 9/7/2026
npx -y skills add modbender/skill-library-mcp --skill local-vosk --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Local Vosk?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/modbender-local-vosk)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: local-vosk
description: Local speech-to-text using Vosk. Lightweight, fast, fully offline. Perfect for transcribing Telegram voice messages, audio files, or any speech-to-text task without cloud APIs.
---
# Local Vosk STT
Lightweight local speech-to-text using Vosk. **Fully offline** after model download.
## Use Cases
- **Telegram voice messages** — transcribe .ogg voice notes automatically
- **Audio files** — any format ffmpeg supports
- **Offline transcription** — no API keys, no cloud, no costs
## Quick Start
```bash
# Transcribe Telegram voice message
./skills/local-vosk/scripts/transcribe voice_message.ogg
# Transcribe any audio
./skills/local-vosk/scripts/transcribe audio.mp3
# With language (default: en-us)
./skills/local-vosk/scripts/transcribe audio.wav --lang en-us
```
## Supported Formats
Any format ffmpeg can decode: **ogg** (Telegram), mp3, wav, m4a, webm, flac, etc.
## Models
Default model: `vosk-model-small-en-us-0.15` (~40MB)
Other models available at https://alphacephei.com/vosk/models
## Setup (if not installed)
```bash
pip3 install vosk --user --break-system-packages
# Download model
mkdir -p ~/vosk-models && cd ~/vosk-models
wget https://alphacephei.com/vosk/models/vosk-model-small-en-us-0.15.zip
unzip vosk-model-small-en-us-0.15.zip
```
## Notes
- Quality is good for conversational speech
- For higher accuracy, use larger models or faster-whisper
- Processes audio at ~10x realtime on typical hardware
- Telegram voice messages are .ogg format — works out of the box
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!