Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Speech Recognition

ASecurity

Transcribe live and recorded audio with the Speech framework. Use when implementing SpeechAnalyzer (iOS 26+), SpeechTranscriber, SFSpeechRecognizer, live microphone speech-to-text, audio file transcription, on-device speech models, speech recognition permissions, or custom language models.

3 stars
0 votes
0 copies
0 views
Added 9/28/2026
developmentswiftnodedocumentation

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add thiennc-tesoglobal/ios-skills --skill speech-recognition --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Speech Recognition?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Speech Recognition
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thiennc-tesoglobal-speech-recognition/badge)](https://www.skillsdirectory.com/skills/thiennc-tesoglobal-speech-recognition)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: speech-recognition
description: "Transcribe live and recorded audio with the Speech framework. Use when implementing SpeechAnalyzer (iOS 26+), SpeechTranscriber, SFSpeechRecognizer, live microphone speech-to-text, audio file transcription, on-device speech models, speech recognition permissions, or custom language models."
---

# Speech Recognition

Convert live speech and prerecorded audio to text. Targets modern on-device `SpeechAnalyzer` (iOS 26+) and legacy `SFSpeechRecognizer` (iOS 10+).

## Contents

- [Permissions & Audio Session](#permissions--audio-session)
- [Modern SpeechAnalyzer (iOS 26+)](#modern-speechanalyzer-ios-26)
- [Legacy SFSpeechRecognizer](#legacy-sfspeechrecognizer)
- [On-Device Model Assets](#on-device-model-assets)
- [Common Mistakes](#common-mistakes)
- [Review Checklist](#review-checklist)
- [References](#references)

## Permissions & Audio Session

Add `NSSpeechRecognitionUsageDescription` and `NSMicrophoneUsageDescription` to Info.plist. Request speech and microphone authorization before starting recording:

```swift
import Speech
import AVFAudio

func requestPermissions() async -> Bool {
    let speechAuthorized = await withCheckedContinuation { continuation in
        SFSpeechRecognizer.requestAuthorization { status in
            continuation.resume(returning: status == .authorized)
        }
    }
    guard speechAuthorized else { return false }
    return await AVAudioApplication.requestRecordPermission()
}
```

## Modern SpeechAnalyzer (iOS 26+)

`SpeechAnalyzer` delivers modular, asynchronous streaming transcription using `SpeechTranscriber`:

```swift
import Speech

final class LiveTranscriber {
    private var analyzer: SpeechAnalyzer?
    private var transcriber: SpeechTranscriber?

    func startTranscribing(format: AVAudioFormat) async throws {
        let transcriber = SpeechTranscriber(locale: Locale(identifier: "en-US"), preset: .transcription)
        let analyzer = SpeechAnalyzer(modules: [transcriber])
        self.transcriber = transcriber
        self.analyzer = analyzer

        Task {
            for try await result in transcriber.results {
                let text = result.text
                if result.isFinal {
                    print("Committed: \(text)")
                } else {
                    print("Partial: \(text)")
                }
            }
        }

        try await analyzer.start(inputFormat: format)
    }

    func appendAudio(buffer: AVAudioPCMBuffer) {
        analyzer?.append(buffer)
    }

    func stop() async throws {
        try await analyzer?.finish()
    }
}
```

## Legacy SFSpeechRecognizer

For audio file transcription and backward compatibility:

```swift
func transcribeFile(url: URL) async throws -> String {
    guard let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US")),
          recognizer.isAvailable else {
        throw SpeechError.unavailable
    }

    let request = SFSpeechURLRecognitionRequest(url: url)
    request.requiresOnDeviceRecognition = true

    return try await withCheckedThrowingContinuation { continuation in
        recognizer.recognitionTask(with: request) { result, error in
            if let error {
                continuation.resume(throwing: error)
            } else if let result, result.isFinal {
                continuation.resume(returning: result.bestTranscription.formattedString)
            }
        }
    }
}
```

## On-Device Model Assets

Check asset installation or download language packs before recognition:

```swift
let status = await AssetInventory.status(for: .transcription, locale: Locale(identifier: "en-US"))
if status != .installed {
    try await AssetInventory.download(for: .transcription, locale: Locale(identifier: "en-US"))
}
```

## Common Mistakes

- **Missing microphone usage description**: Omitting `NSMicrophoneUsageDescription` crashes immediately when accessing `AVAudioEngine`.
- **Conflating partial and final results**: Partial results update frequently; only commit text into data models when `result.isFinal` is true.
- **Starting multiple concurrent recognition tasks**: Always finish or cancel the previous task before launching a new recognition stream.
- **Ignoring on-device fallback failures**: If `requiresOnDeviceRecognition` is true, ensure model assets are installed via `AssetInventory`.
- **Not stopping audio engine alongside recognition**: Failing to halt `AVAudioEngine` keeps the microphone indicator active in the status bar.

## Review Checklist

- [ ] `NSSpeechRecognitionUsageDescription` and `NSMicrophoneUsageDescription` present in Info.plist
- [ ] Speech recognition authorization requested and confirmed authorized
- [ ] Audio engine input nodes and tap callbacks cleaned up upon completion
- [ ] `SpeechAnalyzer` (iOS 26+) used for modern streaming pipelines
- [ ] Transient partial results visually distinguished from finalized transcripts

## References

- SpeechAnalyzer pipelines, volatile results, and audio buffering: [references/speechanalyzer-patterns.md](references/speechanalyzer-patterns.md)
- [Speech framework](https://sosumi.ai/documentation/speech)
- [SpeechAnalyzer](https://sosumi.ai/documentation/speech/speechanalyzer)
- [SpeechTranscriber](https://sosumi.ai/documentation/speech/speechtranscriber)
- [SpeechTranscriber.Preset](https://sosumi.ai/documentation/speech/speechtranscriber/preset)
- [DictationTranscriber](https://sosumi.ai/documentation/speech/dictationtranscriber)
- [SpeechDetector](https://sosumi.ai/documentation/speech/speechdetector)
- [SFSpeechRecognizer](https://sosumi.ai/documentation/speech/sfspeechrecognizer)
- [SFSpeechAudioBufferRecognitionRequest](https://sosumi.ai/documentation/speech/sfspeechaudiobufferrecognitionrequest)
- [SFSpeechURLRecognitionRequest](https://sosumi.ai/documentation/speech/sfspeechurlrecognitionrequest)
- [SFSpeechRecognitionResult](https://sosumi.ai/documentation/speech/sfspeechrecognitionresult)
- [SFSpeechRecognitionRequest](https://sosumi.ai/documentation/speech/sfspeechrecognitionrequest)
- [AssetInventory](https://sosumi.ai/documentation/speech/assetinventory)
- [Asking Permission to Use Speech Recognition](https://sosumi.ai/documentation/speech/asking-permission-to-use-speech-recognition)
- [Recognizing Speech in Live Audio](https://sosumi.ai/documentation/speech/recognizing-speech-in-live-audio)
- [Bring advanced speech-to-text to your app with SpeechAnalyzer](https://sosumi.ai/videos/play/wwdc2025/277)

Attribution

thiennc-tesoglobalthiennc-tesoglobal
View sourceMore from thiennc-tesoglobal →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

284972 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2222 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

10311 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →