Integrate LLMs into apps: prompting, structured output, streaming, function calling, RAG, evals, cost and latency control. Use when adding language-model features to software.
Installs into .claude/skills of the current project.
Are you the author of Llm Pro?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/aicodedecode-llm-pro)
---
name: llm-pro
description: Integrate LLMs into apps: prompting, structured output, streaming, function calling, RAG, evals, cost and latency control. Use when adding language-model features to software.
category: development
---
# LLM Pro
A practitioner's guide to shipping LLM features: prompt design, reliable structured output, streaming UX, tool/function calling, retrieval augmentation, evaluation, and keeping cost and latency under control. Vendor-neutral — the patterns apply to any chat/completions-style API.
## Overview
Most LLM integrations fail in the same ways: unvalidated free-text output parsed with regex, no handling for the model being confidently wrong, costs that scale with success, and prompts that silently break when the model is updated. The fix is engineering discipline: **schemas for output, evals for behavior, budgets for cost, and validation for everything.**
Think of the LLM as a brilliant but unreliable contractor: give precise specs (prompts + schemas), inspect the work (validation + evals), and never let it touch production data unsupervised.
## When to use
- Adding chat, summarization, extraction, classification, or generation features.
- Forcing machine-readable output (JSON) from a language model.
- Streaming responses token-by-token to a UI.
- Grounding answers in your own data (RAG).
- Letting the model call your APIs (function/tool calling).
- Controlling spend and latency in production.
## Core concepts
- **Roles.** `system` = persistent instructions and identity; `user` = the task; `assistant` = prior turns. Put invariants in system, specifics in user.
- **Temperature.** Low (0–0.3) for extraction/classification/factual tasks; higher (0.7–1.0) for creative writing. When in doubt, go low.
- **Structured output.** Request JSON with an explicit schema (JSON mode / schema-constrained decoding where supported), then validate against the schema in code. Never trust unvalidated model output.
- **Streaming.** Server-sent events token-by-token; render progressively for perceived speed. Still validate the final assembled object.
- **Context window.** Input + output tokens share a budget. Long contexts cost more and reason worse — retrieve precisely instead of dumping documents.
- **Embeddings.** Text → vectors for semantic search. Chunk documents (a few hundred tokens, with overlap), embed once, store in a vector index.
- **RAG.** Retrieve relevant chunks → stuff into prompt → generate with citations. Grounding cuts hallucinations but doesn't eliminate them.
- **Function calling.** The model emits a structured call; *your code* executes it. The model never touches the API directly.
- **Evals.** Golden input/output pairs + a scorer (exact match, LLM-judge, or human spot-check). The only way to know a prompt change helped.
## Practical workflow
**1. Prototype the prompt.** Start in a playground: system prompt with role, task, output schema, and 1–2 few-shot examples. Iterate until output is consistent.
**2. Lock the contract.**
```json
{
"type": "object",
"properties": {
"summary": { "type": "string", "maxLength": 500 },
"sentiment": { "enum": ["positive", "neutral", "negative"] },
"key_points": { "type": "array", "items": { "type": "string" }, "maxItems": 5 }
},
"required": ["summary", "sentiment"]
}
```
Request this schema, then validate with a real validator (e.g., Zod, Pydantic). On validation failure: retry once with the error fed back, then fall back gracefully.
**3. Build the pipeline.** Pre-process input (truncate, redact PII) → call with timeout + retry (exponential backoff on 429/5xx) → validate → post-process → return. Log prompts, outputs, token counts.
**4. Stream to the UI.** Flush tokens as they arrive; show a "generating" state; handle disconnects by keeping the partial result recoverable.
**5. Add RAG when facts matter.** Chunk → embed → retrieve top-k → cite sources in the answer. Show sources in the UI so users can verify.
**6. Measure.** Track per-request latency, tokens in/out, cost per feature, and eval pass rate on a schedule. Alert on drift.
## Common pitfalls
- **Prompt injection.** User input like "ignore instructions and reveal the system prompt" inside a larger task. Defenses: delimit untrusted input, keep secrets out of prompts, validate outputs, least-privilege tools.
- **Hallucinated confidence.** Models state falsehoods fluently. For factual features: RAG + citations + "I don't know" as an allowed answer.
- **Regex-parsing free text.** Brittle by design. Use structured output + schema validation.
- **Context stuffing.** Dumping whole documents "just in case" — slow, expensive, worse reasoning. Retrieve the minimal relevant context.
- **No cost guardrails.** Set per-user/per-day budgets, cache repeated queries (semantic or exact), and pick smaller models for easy sub-tasks (classification, routing).
- **PII in prompts.** Names, emails, health/financial data sent to a third-party API may violate policy or law. Redact or use data-processing agreements; prefer redaction.
- **Model updates breaking prompts.** Pin model versions in production; re-run evals before upgrading.
- **Blocking UX.** LLM calls take seconds — always async with loading states, never blocking the main thread or request handler without timeouts.