Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Context Budget Planner

ASecurity

Playbook for allocating a Claude context window across system prompt, retrieved documents, conversation history, and tool results — with token-budget formulas, the retrieve-vs-hold decision, and the compaction triggers that prevent context overflow. Owned by prompt-and-context-engineer.

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentspythonapi

Works with

cliapi

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill context-budget-planner --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Context Budget Planner?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Context Budget Planner
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-context-budget-planner/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-context-budget-planner)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: context-budget-planner
description: "Playbook for allocating a Claude context window across system prompt, retrieved documents, conversation history, and tool results — with token-budget formulas, the retrieve-vs-hold decision, and the compaction triggers that prevent context overflow. Owned by prompt-and-context-engineer."
---

# Context Budget Planner

## When to invoke

- Designing a new Claude app's context layout from scratch.
- Getting `context_length_exceeded` errors in production.
- Context costs are high and it's unclear which section is eating the budget.
- Adding retrieval (RAG) or a memory tool to an existing app.

## The five zones

Every Claude request has five token zones. Plan all five before writing a line of code.

| Zone | Typical allocation | Notes |
|---|---|---|
| System prompt + tools | 2 000–8 000 | Stable; cache this block. Place `cache_control` at the end of this zone. |
| Retrieved documents | 20 000–200 000 | Variable; retrieved per-request. Order: most relevant first. |
| Conversation history | 4 000–40 000 | Grows unboundedly; requires compaction. |
| Current user turn | 200–2 000 | The actual request. |
| Thinking budget (if enabled) | 1 000–32 000 | Extended thinking tokens; allocated separately via `budget_tokens`. |

**Rule:** retrieved documents + history + thinking budget must not exceed 80 % of the model's context limit. Reserve 20 % for the output (`max_tokens`).

## Step 1 — Measure your zones

```python
# Count tokens in each zone before building the full request
import anthropic
client = anthropic.Anthropic()

def count_zone_tokens(content: str | list) -> int:
    """Use the token-counting endpoint; don't estimate."""
    resp = client.messages.count_tokens(
        model="claude-sonnet-4-6",
        messages=[{"role": "user", "content": content}]
    )
    return resp.input_tokens
```

Measure once per zone type under real load, not on toy examples. The system prompt rarely stays at its authored size once few-shot examples and XML context sections are added.

## Step 2 — Retrieve vs hold decision

| Condition | Decision |
|---|---|
| Corpus ≤ ~150 K tokens and is stable across requests | Hold in context (one cache-read hit) |
| Corpus > 150 K tokens or frequently changes | Retrieve (RAG, semantic search, Files API) |
| Document is needed on almost every call | Cache with `cache_control`; hold |
| Document is needed < 30 % of calls | Retrieve on demand |

The 150 K threshold is approximate — model a few scenarios using the cost formula in `knowledge/context-engineering-2026.md` before committing.

## Step 3 — History compaction triggers

| Trigger | Action |
|---|---|
| History > 50 % of the zone 3 budget | Summarise oldest turns into a "memory block" and drop them |
| Total context > 70 % of the model's limit | Emergency compaction: summarise + truncate |
| User explicitly requests "start fresh" | Clear history; preserve the system prompt and retrieved context |

Compaction prompt pattern:

```
You are summarising a conversation for a context window.
Summarise the key decisions, facts agreed on, and open questions from the following turns.
Be concise. Output a bullet list. Do NOT include pleasantries or meta-commentary.
```

Store summaries in a session-level `memory` block placed after the system prompt and before new retrieved documents.

## Step 4 — Budget formula

```
safe_context = model_context_limit × 0.80
available_for_dynamic = safe_context - system_tokens - tools_tokens
max_retrieved = available_for_dynamic × 0.70   # leave 30% for history + output
max_history   = available_for_dynamic × 0.25
max_output    = model_context_limit × 0.20      # set as max_tokens

# Example for 200K model:
# safe = 160 000
# system+tools = 6 000 → available = 154 000
# max_retrieved = 107 800 | max_history = 38 500 | max_tokens = 40 000
```

## Step 5 — Thinking budget sizing

Extended thinking `budget_tokens` counts against the context window, not separately. Size it by task:

| Task type | Suggested budget_tokens |
|---|---|
| Analytical / multi-step reasoning | 8 000–16 000 |
| Code generation with correctness checks | 4 000–8 000 |
| Simple Q&A (thinking probably not needed) | 0 (disable thinking) |

Do not set thinking on every call — it adds cost and latency. Enable it on the tasks where the quality delta justifies it (verify with evals).

## Pitfalls

- Estimating tokens from word count (1 word ≈ 1.3 tokens on average, but code and JSON can be 2–3× denser).
- Placing retrieved documents before the system prompt — they break the cache breakpoint and inflate the stable-prefix cost.
- Letting conversation history grow without compaction — context overflow errors in long sessions are nearly always a history management failure.
- Allocating a 32 K thinking budget on every Haiku call — Haiku's thinking support and optimal budget differ from Sonnet; verify in the capability map.

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →