Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Vision Framework

ASecurity

Builds or reviews iOS computer-vision features with Vision and VisionKit, including OCR, barcode and document scanning, face/object detection, segmentation, tracking, and Core ML inference. Use for Vision requests, live DataScanner flows, or Vision/Core ML integration.

3 stars
0 votes
0 copies
0 views
Added 9/28/2026
developmentswiftexpressapidocumentation

Works with

api

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add thiennc-tesoglobal/ios-skills --skill vision-framework --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Vision Framework?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Vision Framework
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thiennc-tesoglobal-vision-framework/badge)](https://www.skillsdirectory.com/skills/thiennc-tesoglobal-vision-framework)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: vision-framework
description: "Builds or reviews iOS computer-vision features with Vision and VisionKit, including OCR, barcode and document scanning, face/object detection, segmentation, tracking, and Core ML inference. Use for Vision requests, live DataScanner flows, or Vision/Core ML integration."
---

# Vision Framework

Detect text, faces, barcodes, objects, contours, and poses in images and live video using Apple's on-device computer vision frameworks (`Vision` and `VisionKit`). Targets Swift 6.3 / iOS 26+.

## Contents

- [Two API Generations](#two-api-generations)
- [Modern Request Architecture](#modern-request-architecture)
- [Normalized Coordinates vs Pixel Space](#normalized-coordinates-vs-pixel-space)
- [Core ML Integration](#core-ml-integration)
- [Route by Task](#route-by-task)
- [Common Mistakes](#common-mistakes)
- [Review Checklist](#review-checklist)
- [References](#references)

## Two API Generations

| Dimension | Modern Vision (iOS 18+) | Legacy Vision (pre-iOS 18) |
|---|---|---|
| Request Types | Swift structs (`RecognizeTextRequest`, `DetectBarcodesRequest`) | ObjC classes (`VNRecognizeTextRequest`, `VNDetectBarcodesRequest`) |
| Execution | `try await request.perform(on: image)` | `VNImageRequestHandler(cgImage:).perform([request])` |
| Results | Strongly typed observation arrays | `request.results as? [VNBarcodeObservation]` |
| Concurrency | Native async/await | Completion handlers or synchronous blocking calls |

Prefer modern Swift-native request structs for new code. Maintain legacy handlers only when supporting deployment targets below iOS 18.

## Modern Request Architecture

All modern requests conform to `ImageProcessingRequest`:
```swift
import Vision

var request = RecognizeTextRequest()
request.recognitionLevel = .accurate
request.recognitionLanguages = [Locale.Language(identifier: "en-US")]

let observations = try await request.perform(on: cgImage)
for observation in observations {
    print("Found text: \(observation.topCandidates(1).first?.string ?? "")")
}
```

## Normalized Coordinates vs Pixel Space

Vision observations express bounding boxes in normalized coordinates `(0.0...1.0)` with origin at the **bottom-left** corner (Cartesian), while UIKit/SwiftUI places origin at the **top-left**.

To convert to UIKit/SwiftUI coordinates:
```swift
let rect = VNImageRectForNormalizedRect(observation.boundingBox, Int(viewWidth), Int(viewHeight))
// Flip Y-axis: y = viewHeight - rect.origin.y - rect.size.height
```

## Core ML Integration

Run custom Core ML vision models using `CoreMLModelContainer`:
1. Compile model (`.mlpackage` or `.mlmodelc`).
2. Wrap with `CoreMLModelContainer(model: compiledModel)`.
3. Create image request: `var request = ImageFeaturePrintRequest()` or custom Core ML classification request.

## Route by Task

- For OCR text recognition, barcode scanning, face detection, person segmentation, and object tracking, read [Vision Requests and Detectors](references/vision-requests.md).
- For live camera scanner UI, document scanning, and optical barcode flows with `DataScannerViewController`, read [VisionKit Scanner](references/visionkit-scanner.md).

## Common Mistakes

- Forgetting to invert the Y-axis when projecting normalized Vision bounding boxes onto UIKit/SwiftUI views.
- Running heavy Vision requests (`.accurate` text recognition) synchronously on the main thread.
- Passing `UIImage` directly without extracting `cgImage` or preserving image orientation metadata.
- Using `DataScannerViewController` without checking `DataScannerViewController.isSupported` and `isAvailable`.
- Retaining stateful tracking requests (`TrackObjectRequest`) across unrelated image sequences.

## Review Checklist

- [ ] Modern request types (`RecognizeTextRequest`) preferred on iOS 18+ targets
- [ ] Image orientation properly passed to `perform(on:orientation:)`
- [ ] Vision bounding box coordinates converted and Y-flipped for SwiftUI rendering
- [ ] Camera scanner checks `DataScannerViewController.isSupported` before presentation
- [ ] Heavy processing executed on background tasks off `@MainActor`
- [ ] Bounding boxes clamped to image bounds before cropping

## References

- [Vision requests, legacy VNRequest patterns, and Core ML integration](references/vision-requests.md)
- [VisionKit DataScannerViewController and document scanner](references/visionkit-scanner.md)
- [Vision documentation](https://sosumi.ai/documentation/vision)
- [VisionKit documentation](https://sosumi.ai/documentation/visionkit)
- [RecognizeTextRequest](https://sosumi.ai/documentation/vision/recognizetextrequest)

Attribution

thiennc-tesoglobalthiennc-tesoglobal
View sourceMore from thiennc-tesoglobal →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

284972 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2222 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

10311 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →