Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Instructor Library

ASecurity

Get structured, validated outputs from LLMs with Instructor — Pydantic schemas, retries, and reliable JSON extraction.

2 stars
0 votes
0 copies
5 views
Added 9/29/2026
ai-agentspythonrustgoapidatabase

Works with

cliapi

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill instructor-library --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Instructor Library?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Instructor Library
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-instructor-library/badge)](https://www.skillsdirectory.com/skills/aicodedecode-instructor-library)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: instructor-library
description: Get structured, validated outputs from LLMs with Instructor — Pydantic schemas, retries, and reliable JSON extraction.
category: ai-research
---

## Overview

Instructor is a Python library that turns LLM outputs into validated structured
data. You define a Pydantic model describing exactly the shape you want;
Instructor patches your LLM client so responses are parsed and validated against
that schema, with automatic retries when validation fails. Instead of
prompt-engineering JSON out of a model and praying the parse holds, you declare
the schema and let the library enforce it.

The pattern is simple and robust: schema as contract, validation as gate, retries
as recovery. Instructor supports multiple providers and modes (function calling,
JSON mode, markdown code blocks), picking the most reliable mechanism per
provider. It's the default answer to "I need the model to return structured data
I can actually use in code."

Instructor's philosophy: the schema is the prompt. A well-described Pydantic
model teaches the model what you want better than paragraphs of instructions, and
the validator catches what the prompt misses.

## When to use

- Extracting structured data from text: entities, attributes, classifications with
  fields.
- Any LLM-to-code boundary where downstream code consumes the output (APIs,
  databases, pipelines).
- Classification tasks where you want labels plus structured rationales or
  confidences.
- Multi-field generation: summaries with metadata, structured plans, form
  filling.
- Replacing fragile regex/JSON-parse postprocessing with validated schemas.
- Batch extraction jobs where per-item reliability compounds.

## Core concepts

- **Response models**: Pydantic classes defining the output schema. Field types,
  descriptions, and constraints (min/max, patterns, enums) all become instructions
  to the model and validation rules on the output. Descriptions are prompt
  engineering — write them for the model.
- **Modes**: function-calling mode (most reliable where supported), JSON mode,
  and fallback parsing modes. Instructor abstracts the provider differences;
  choose the strictest mode your provider supports.
- **Automatic retries**: when validation fails, Instructor re-asks with the error
  message included, up to a configured limit. This converts most malformed outputs
  into valid ones without your code handling it.
- **Streaming and partials**: for long outputs, stream partial objects as they're
  generated. Useful for UX (progressive display) and for failing fast on doomed
  generations.
- **Validation as prompt**: Pydantic validators (e.g., "end_date must be after
  start_date") run on model output and their error messages go back to the model
  on retry — the model learns the constraint from its own failure.
- **Batch and async**: first-class async support for high-throughput extraction.
  Structure your extraction as concurrent validated calls, not sequential hope.
- **Field descriptions as instructions**: each description should say what the
  field means, give an example value, and state constraints. This is the
  highest-leverage prompt content in structured extraction.
- **Nested models**: compose complex schemas from smaller Pydantic models. Keeps
  schemas readable and lets you reuse sub-schemas across tasks.

## Practical workflow

1. **Design the schema.** Write the Pydantic model: fields, types, descriptions,
   constraints. Keep it as small as the use case allows — every field is a chance
   for the model to err.
2. **Write field descriptions for the model.** Each description should say what
   the field means and give an example value. This is the highest-leverage prompt
   content in structured extraction.
3. **Pick the mode.** Function-calling where available; JSON mode otherwise. Test
   that the mode is actually supported by your provider+model combo.
4. **Set retries sensibly.** 2–3 retries with the validation error fed back. More
   retries rarely fix systematic schema misunderstandings — those need schema or
   prompt fixes.
5. **Handle the residual failures.** Even with retries, some outputs won't
   validate. Decide: fall back to a default, escalate to human review, or fail
   loudly. Never silently accept unvalidated data.
6. **Evaluate extraction quality.** On a labeled sample: field-level accuracy, not
   just "valid JSON." A schema-valid output with wrong values is worse than a
   parse error — it's silently wrong.
7. **Load-test the pattern.** Structured extraction at volume: measure latency,
   retry rates, and cost per successful extraction. High retry rates signal schema
   problems, not bad luck.

Checklist for production structured output:
- Schema minimal and precisely described per field.
- Retry count set; residual-failure policy defined.
- Field-level accuracy measured on labeled data.
- Unvalidated outputs never flow silently downstream.
- Retry rate monitored (spikes indicate schema drift).

## Common pitfalls

- **Schemas too ambitious.** Asking for 30 fields in one call degrades every
  field's accuracy. Split into multiple focused extractions.
- **Vague field descriptions.** "summary: str" gets you whatever the model feels
  like. "summary: one sentence, no quotes, under 200 characters" gets you that.
- **Trusting validity as correctness.** Validation checks shape, not truth. A
  valid-but-hallucinated extraction needs content-level evaluation.
- **Retry storms.** High retry counts on a systematically misunderstood schema
  burn money and latency. If retry 1 usually fails, fix the schema.
- **No fallback policy.** The 1% that never validates will happen in production.
  Decide up front what happens — don't let it throw in the request path unhandled.
- **Provider mode mismatch.** Assuming function calling works identically
  everywhere. Verify per provider; behavior differs in edge cases.
- **Nested schema overload.** Deeply nested schemas with dozens of fields in one
  call. Flatten or split — the model's accuracy degrades with schema complexity.
- **Ignoring retry-rate signals.** A rising retry rate in production means the
  model, the data, or the schema changed. Alert on it; don't just absorb the
  cost.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →