Skip to content
Back to skills

Audio Transcribe

ASecurity

Transcribe speech from audio files (mp3, m4a, wav, ogg, flac, webm) to text using the local `whisper` CLI — no API key. Use whenever a task hinges on the spoken content of an audio attachment.

  • 17 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 4, 2026
toolsgoshellgitapi

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 4, 2026

npx -y skills add gabrielmoreira/agent-skills-mirror --skill audio-transcribe --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Audio Transcribe?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Audio Transcribe
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/gabrielmoreira-audio-transcribe/badge)](https://www.skillsdirectory.com/skills/gabrielmoreira-audio-transcribe)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: audio-transcribe
description: Transcribe speech from audio files (mp3, m4a, wav, ogg, flac, webm) to text using the local `whisper` CLI — no API key. Use whenever a task hinges on the spoken content of an audio attachment.
version: 1.0.0
requires_tools:
  - os.shell.run
  - os.fs.read
dangerous: false
platforms:
  - darwin
  - linux
---

# audio-transcribe

Turn spoken audio into text locally with [`whisper`](https://github.com/openai/whisper).
Whisper writes a plain-text transcript next to the audio; you then read that file
and answer from its content. The chat model never "hears" the audio — `whisper`
does the listening as a separate process, so this works regardless of model size.

## Setup health check (run first, every session)

Verify with **one solo step**:

```
[{ "tool": "os.shell.run", "args": { "cmd": "whisper", "args": ["--help"] } }]
```

Outcome map:
- `exit 0` + usage text → ready, proceed.
- `command not found: whisper` → enter **Setup playbook → "whisper missing"**.

## Setup playbook (when prerequisites are missing)

### whisper missing

Reply (solo `reply` step):

> "`whisper` is not installed. I can install it: `brew install openai-whisper` (also needs `ffmpeg`). Install it?"

On yes:

```
[{ "tool": "os.shell.run", "args": { "cmd": "brew", "args": ["install", "openai-whisper", "ffmpeg"] } }]
```

On Linux: `pipx install openai-whisper` (or `pip install openai-whisper`) plus
`apt-get install ffmpeg`. The first transcription downloads the model weights to
`~/.cache/whisper`.

## When to use

- The task references an audio attachment (`.mp3`, `.m4a`, `.wav`, `.ogg`,
  `.flac`, `.webm`) and the answer depends on what is said in it.
- "What does the speaker say…", "list the ingredients mentioned…", "which page
  numbers are read aloud…".

## When NOT to use

- The audio only needs format conversion / trimming — that's the `ffmpeg` skill.
- The attachment is an image / document — use vision or `fs.read_document`.

## How to transcribe

1. Transcribe to a `.txt` next to the audio (one solo step). Pick the model by
   need: `small` is a good speed/accuracy default; use `medium` when accuracy
   matters and the clip is short.

```
[{ "tool": "os.shell.run", "args": { "cmd": "whisper", "args": ["audio.mp3", "--model", "small", "--output_format", "txt", "--output_dir", ".", "--language", "en"] } }]
```

2. Read the produced transcript (whisper names it `<basename>.txt`):

```
[{ "tool": "os.fs.read", "args": { "path": "audio.txt" } }]
```

3. Answer strictly from the transcript text. If a list/order is requested,
   preserve the spoken order exactly.

## Rules

1. Always write the transcript to a NEW `.txt`; never overwrite the source audio.
2. Drop `--language` only when the spoken language is unknown; setting it
   correctly improves accuracy and speed.
3. Long clips re-encode slowly — set realistic expectations and prefer a smaller
   `--model` for multi-minute audio.
4. Base the final answer on the transcript content, not on the file name or
   metadata.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…