Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Rag Pro

ASecurity

RAG guidance — chunking, retrieval tuning, hybrid search, reranking, citations, evaluation, and production RAG.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongo

Works with

cli

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill rag-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Rag Pro?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Rag Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-rag-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-rag-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: rag-pro
description: RAG guidance — chunking, retrieval tuning, hybrid search, reranking, citations, evaluation, and production RAG.
category: development
---

## Overview

RAG (Retrieval-Augmented Generation) grounds LLMs in your data: retrieve relevant chunks, stuff them into the prompt, generate answers with citations. It's the highest-ROI LLM pattern — and the one where "it works in the demo" most often collapses in production. The unglamorous truth: chunking, retrieval quality, and evaluation matter far more than model choice.

This skill covers production RAG end to end: ingestion and chunking, retrieval tuning (hybrid, reranking, filters), synthesis with citations, evaluation, and the operational concerns of a live system.

## When to use

- Building question-answering over documents.
- Improving RAG answer quality.
- Tuning chunking and retrieval.
- Adding citations to generated answers.
- Evaluating RAG systems.
- Operating RAG in production (updates, access control, monitoring).

## Core concepts

- **The pipeline.** Ingest → chunk → embed → index → retrieve → rerank → synthesize → cite. Every stage has quality levers; failures compound — debug stage by stage, not end-to-end.
- **Chunking.** The highest-leverage decision: semantic boundaries (sections, paragraphs) over fixed sizes; 300-800 tokens typical with overlap; hierarchical (small chunks + parent context); metadata per chunk (source, section, date). Bad chunks cap everything downstream.
- **Embeddings.** Domain-appropriate model, fixed for the index; chunk text + metadata strategies (what gets embedded vs filtered). Re-embed everything on model change.
- **Retrieval.** Top-k similarity as the baseline; `similarity_top_k` tuning (more ≠ better — noise dilutes); metadata pre-filtering (tenant, date, source) before similarity.
- **Hybrid search.** Dense + BM25 fused (RRF) — semantic for paraphrases, keyword for exact terms (names, codes, SKUs). Usually the single biggest retrieval win after chunking.
- **Reranking.** Cross-encoder or LLM rerank over top-20→top-5 — the cheapest quality improvement in the pipeline; slower but applied to few candidates.
- **Query transformations.** HyDE (hypothetical answer embedding), query rewriting, multi-query expansion — bridging the vocabulary gap between questions and documents. Useful when queries and docs speak different languages.
- **Context assembly.** Order (best first? best last — "lost in the middle"), deduplication, token budgeting, `LongContextReorder`. More context isn't better — relevant context is.
- **Synthesis.** The generation prompt: "answer ONLY from context," citation requirements, refusal behavior for unanswerable questions. The prompt is a reliability control, not decoration.
- **Citations.** Chunk IDs → source references in answers — verifiability is the feature; users must be able to check. Citation format designed for the UI (inline, footnotes, side panel).
- **Refusal.** "I don't know" as a designed behavior — unanswerable questions answered confidently are the worst failure mode. Thresholds on retrieval scores + explicit instructions.
- **Evaluation.** Two layers: retrieval (hit rate, MRR@k on labeled Q&A) and answer (faithfulness — grounded in context? relevancy — answers the question?). RAGAS-style metrics automate the answer layer; human review on samples.
- **Incremental updates.** Document changes → re-chunk/re-embed affected docs → upsert; deletions as tombstones; versioned corpora. Stale RAG lies with confidence.
- **Access control.** Metadata-based filtering (user's clearance ≥ chunk's classification) — enforced at retrieval, tested adversarially. RAG over mixed-sensitivity corpora without ACLs is a breach waiting.
- **Monitoring.** Retrieval scores distribution, refusal rates, citation click-through, user feedback (thumbs up/down per answer) — quality signals from production, feeding the eval set.
- **Graph RAG.** Knowledge-graph-augmented retrieval for multi-hop questions — entities and relationships as the retrieval structure; powerful for connected data, heavy to build.
- **Agentic RAG.** The retriever as an agent tool — the LLM decides what to search, iteratively; flexible for complex research, costlier and less predictable than fixed pipelines.

## Practical workflow

1. **Build the eval set first.** 50-100 real questions with known-good answers and source chunks — before tuning anything. This set judges every later decision.
2. **Ingest and chunk deliberately.** Semantic boundaries, overlap, rich metadata; inspect chunks manually — read 50 random chunks before indexing:
   ```python
   chunks = semantic_chunk(documents, target_tokens=512, overlap=50,
                           metadata=["source", "section", "date"])
   ```
3. **Baseline retrieval.** Dense top-k; measure hit rate/MRR on the eval set. This number anchors all improvements.
   ```python
   # retrieval eval: hit rate on labeled Q&A
   hits = sum(gold_chunk in [c.id for c in retrieve(q, k=5)]
              for q, gold_chunk in eval_set)
   print(f"hit@5: {hits / len(eval_set):.2f}")
   ```

4. **Add hybrid + rerank.** BM25 fusion (RRF), then cross-encoder rerank top-20→5; measure each addition independently — keep what moves the metric.
5. **Tune assembly.** Token budgets, ordering, dedup; verify the context actually fits and the best chunks aren't buried.
6. **Write the synthesis prompt.** Grounded-only instructions, citation format, refusal behavior:
   ```
   Answer ONLY from the provided context. Cite sources as [1], [2].
   If the context doesn't contain the answer, say "I don't have information on this."
   Never invent details not present in the context.
   ```
7. **Evaluate both layers.** Retrieval metrics + faithfulness/relevancy; human spot-checks on failures; grow the eval set from production misses.
8. **Operate.** Incremental update pipelines, ACL filtering, monitoring (refusal rate, feedback, drift), scheduled re-evaluation — RAG is a living system.

## Common pitfalls

- **Blaming the LLM** — generation is rarely the problem; chunking and retrieval are.
- **No eval set** — tuning by vibes; labeled Q&A before any optimization.
- **Arbitrary chunking** — fixed 1000-char splits mid-sentence; semantic boundaries.
- **Pure dense retrieval** — missing exact terms; hybrid search.
- **Skipping reranking** — cheapest win ignored; rerank top candidates.
- **Context stuffing** — 50 chunks hoping for coverage; top-n discipline.
- **No citations** — unverifiable answers; cite by design.
- **Confident hallucination on gaps** — no refusal behavior; design "I don't know."
- **Stale indexes** — updated docs, old chunks; incremental update pipelines.
- **No access control** — sensitive chunks to everyone; metadata ACLs, tested.
- **Evaluating end-to-end only** — can't separate retrieval vs synthesis failures; two layers.
- **Production without monitoring** — quality decay invisible; feedback loops + re-evaluation.
- **Ignoring query-document mismatch** — jargon gaps; query rewriting/HyDE where needed.
- **Agentic RAG by default** — unpredictable cost/latency; fixed pipelines first, agentic where complexity demands it.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →