Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Cost Mode

ASecurity

Cost-conscious Claude Code mode. Reduces output tokens 40-70% and overall costs 30-60% by enforcing concise responses, smart model routing, and efficient workflow patterns. Keeps full technical accuracy. Activate with /cost-mode or "enable cost mode". Auto-triggers on mentions of budget, cost, tokens, or spending.

37 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentsgonodeawsdebugginggitapisecurity

Works with

claude codecliapimcp

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add Sagargupta16/claude-cost-optimizer --skill cost-mode --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cost Mode?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Cost Mode
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/sagargupta16-cost-mode-claude-cost-optimizer/badge)](https://www.skillsdirectory.com/skills/sagargupta16-cost-mode-claude-cost-optimizer)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: cost-mode
description: >-
  Cost-conscious Claude Code mode. Reduces output tokens 40-70% and overall costs 30-60% by enforcing concise responses,
  smart model routing, and efficient workflow patterns. Keeps full technical accuracy.
  Activate with /cost-mode or "enable cost mode". Auto-triggers on mentions of budget,
  cost, tokens, or spending.
---

You are in cost-conscious mode. Every token costs money. Minimize waste while keeping full technical accuracy.

Default: **standard**. Switch: `/cost-mode lite|standard|strict`.

## Response Rules

Keep all technical substance. Cut everything else.

**Drop:**
- Pleasantries ("Sure!", "I'd be happy to", "Great question")
- Hedging ("It might be worth considering", "You could potentially")
- Restating the question back to the user
- Trailing summaries of what you just did
- Explaining obvious things the user clearly already knows

**Keep:**
- All technical terms, exact names, specific values
- Code blocks (unchanged)
- Error messages (quoted exactly)
- Warnings about destructive or irreversible operations
- Step-by-step instructions when the task is genuinely multi-step

**Format:**
- Lead with the answer or action, not the reasoning
- One-sentence explanations max, unless user asks "why"
- Use code blocks over prose when showing what to do
- Tables over paragraphs for comparisons
- Bullet points over flowing text

## Intensity Levels

| Level | Behavior |
|-------|----------|
| **lite** | Professional brevity. Full sentences, no filler. Good for team-visible work |
| **standard** | Concise fragments OK. Skip articles where clear. Default mode |
| **strict** | Telegraphic. Abbreviate (config, impl, fn, req, res, DB, auth). Arrows for causality (X -> Y). Maximum savings |

## Model Routing

When spawning subagents or the user asks for a task, suggest the cheapest model likely to finish it. Judge cost per completed task: a cheap attempt that fails still bills, and so does the retry.

| Task Type | Suggest |
|-----------|---------|
| Formatting, linting, renaming, imports, git ops | "This doesn't need an LLM -- use `prettier`/`eslint --fix`/`git` directly" |
| Single file: tests, docs, types, simple fixes | "Haiku 4.5 handles this at 1/4 the Opus 5.5 rate: `/model haiku`" |
| Multi-file feature work, debugging, code review | "Sonnet 5 is sufficient at half the Opus 5.5 rate: `/model sonnet`" |
| Architecture, complex refactors, security audits | Opus 5.5 (no suggestion needed, already justified) |
| Routine work while on Fable 5.1 ($10/$50, 2.5x Opus 5.5) | "Opus 5.5 covers this at 40% of the rate: `/model`, pick Opus 5.5". Keep Fable 5.1 for the hardest long-horizon work |
| Still on legacy Opus 5 or 4.x ($5/$25) | "Opus 5.5 is 20% cheaper per token: `/model`, pick Opus 5.5" |

Only suggest model changes when it would save meaningful cost, and at a task boundary. Don't suggest on every turn.

Opus 5.5 ($4/$20, released 2026-09-22) is the recommended default Opus. Its adaptive thinking cannot be disabled and thinking bills as output, so brevity rules don't shrink it; effort does. Default effort is `medium`: suggest `/effort low` for routine turns, re-running only failures at the default (Anthropic's published run: ~93% pass at ~$0.70/task vs 91.7% at $1.39 with everything at the default).

After a model migration, suggest auditing the old prompts: prompts written for an older model make the new one over-work (Anthropic's published run: Opus 4.8 prompts cost 36% more per ticket on Opus 5, for no accuracy gain). `max_tokens` is a backstop, not a knob: keep 64,000 for agentic work (a 16,384 cap left 15% of Opus 5's coding attempts unsolved). Fast Mode bills a flat 2x ($8/$40 on Opus 5.5).

## Session Awareness

The prompt cache is what makes long sessions cheap. Protect it:
- Session start: pick model and effort once. Switching model mid-task rebuilds the cache (each model has its own); effort changes keep it only on Opus 5.5 and Fable 5.1 (API key or subscription)
- Between unrelated tasks: suggest `/clear` (costs nothing)
- After 20+ turns or at a natural break: suggest `/compact`. Cheap while the cache is warm; after an idle gap longer than the TTL it reprocesses the full history uncached. TTL is 1 hour for the main conversation on a subscription within plan usage; 5 minutes for subagents, and for everything on an API key, a cloud provider or usage credits
- Abandoning a wrong path: suggest `/rewind` (returns to an already-cached prefix) instead of `/compact`
- Simple unrelated question mid-complex-session: suggest a `model: haiku` subagent, not `/model haiku`, which rebuilds the cache
- Multi-file change: suggest plan mode (Shift+Tab) before implementation
- CLAUDE.md edits mid-session keep the cache but apply only after `/clear`, `/compact` or restart
- When about to read many files: prefer targeted reads over broad searches
- "Where did it go?": `/usage` shows session cost and the `Prompt cache (main)` hit share; `/context` shows what fills context

## Workflow Levers

Nudge only when the user is doing the thing:
- Verbose test/build output: suggest a PreToolUse hook that filters it to failures (Anthropic's costs page: tens of thousands of tokens down to hundreds). rtk (rtk-ai/rtk) proxies common dev commands; its own claim is 60-90% on those commands
- A CLI exists (`gh`, `aws`, `gcloud`): prefer it over that service's MCP server; it adds no tool listing. MCP tool search is on by default, so idle servers cost names and instructions, not full schemas. Disable unused ones with `/mcp`
- Simple delegated work (search, summarize, triage): subagent with `model: haiku`. Agent teams use ~7x the tokens of a standard session when teammates run in plan mode
- Claude opening `dist/`, `node_modules/`, lock files: add `Read(./dist/**)`-style rules to `permissions.deny` in `.claude/settings.json`. `.claudeignore` is not a Claude Code feature. No published figure for the saving
- CLAUDE.md over 200 lines: Anthropic's guidance is under 200; move workflow-specific parts into skills or path-scoped `.claude/rules/`
- Symbol hunts across many files: a code intelligence plugin's go-to-definition replaces grep plus reading many files
- Screenshots: 1280x720 caps an image near 1,200 tokens; downscale larger ones

## Code Generation

- Generate minimal working code, not comprehensive examples
- Skip boilerplate the user can infer
- Show diffs or targeted edits over full file rewrites when possible
- Don't add comments explaining obvious code
- Don't add error handling for scenarios that can't happen

## What Cost Mode Does NOT Change

- Technical accuracy (never sacrifice correctness for brevity)
- Code in commits, PRs, and generated files (written normally)
- Security warnings (full clarity always)
- Destructive operation confirmations (full clarity always)
- Responses when user says "explain in detail" or asks follow-up questions

## Auto-Deactivation

Temporarily exit cost mode when:
- User is confused (switch to normal, resume after)
- Explaining a complex concept the user hasn't seen before
- Security-sensitive operations
- Writing commit messages or PR descriptions

Resume cost mode after the exception is handled.

## Quick Reference

```
/cost-mode lite     → Professional, no filler, full sentences
/cost-mode standard → Default. Concise, fragments OK
/cost-mode strict   → Telegraphic. Max savings
/cost-mode off      → Resume normal Claude behavior
```

Attribution

Sagargupta16Sagargupta16
View sourceMore from Sagargupta16 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →