Writes an Agent Spec with the goal and done-state, allowed tools and permissions, stop conditions and a spend budget per session, the approval points carried over from the human-in-the-loop design, and the list of actions the agent must never take alone. Use for "run aipm-agent-spec", "agent spec", "write the limits for our agent", "agent permissions", "where should the agent stop", "agent guardrails", "excessive agency", "spend cap for the agent", part of the Claude for AI Product Managers P...
Installs into .claude/skills of the current project.
Are you the author of Aipm Agent Spec?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/polar-bear-org-aipm-agent-spec)
---
name: aipm-agent-spec
description: Writes an Agent Spec with the goal and done-state, allowed tools and permissions, stop conditions and a spend budget per session, the approval points carried over from the human-in-the-loop design, and the list of actions the agent must never take alone. Use for "run aipm-agent-spec", "agent spec", "write the limits for our agent", "agent permissions", "where should the agent stop", "agent guardrails", "excessive agency", "spend cap for the agent", part of the Claude for AI Product Managers Pack by Polar Bear.
---
# Agent Spec
## When To Use
The agent can act on its own and nobody has written where it must stop. It can read the inbox, update records and send replies, and the only limit is a sentence in the prompt. It answers: what may the agent do, with which permissions, when does it stop, and what may it never finish without a person?
## When Not To Use
If nobody has decided that an agent is needed, run the Workflow or Agent Decision first: a fixed chain needs no agent spec. If the approval points are not designed yet, run Human-in-the-Loop Design first; this spec copies them, it does not invent them.
## Inputs
- The goal of the agent, the tools it could call and the systems behind them, with today's access rights
- Your Human-in-the-Loop Design, and any transcripts of the agent running, with personal data masked
If you have none of this, I start from the goal and a list of actions it would take, and mark the output as a first draft.
## Approach
The structure comes from two public sources. Anthropic Engineering's "Building effective agents" sets the frame: agents trade cost and latency for autonomy, so they need stopping conditions and human checkpoints. The OWASP Top 10 for LLM Applications 2025, entry LLM06 Excessive Agency, names the three root causes to close: too much functionality, too many permissions, too much autonomy. The failure it prevents: an agent asked to tidy a shared folder that had delete rights, read an instruction hidden in a document and did exactly what it said.
## Workflow
1. Ask three questions: what does finished look like, which systems can it touch, and who signs the budget and the never-alone list?
2. Write the goal and the done-state as a checkable final state in the system, not the agent's own "done" message.
3. Fill the three excessive agency tables. Functionality: only the tools the goal needs, narrow tools over open-ended ones such as a shell. Permissions: least privilege, read-only wherever possible, scoped to the user. Autonomy: which actions need a person.
4. Set the stop conditions: maximum steps, time, spend budget per session, repeated failure on the same step, and a blocker reached. Each ends in "stop and hand to a person". The hard cap is a blank for a named person. On Claude Managed Agents (beta), session spend budgets and permission policies can enforce these; on any other stack, engineering builds the equivalent.
5. Copy the approval points from the Human-in-the-Loop Design with their triggers. Do not redesign them here; if one is missing, send it back to that design.
6. Write the never-alone list: sending, paying, deleting, changing permissions, contacting customers, and anything else the user adds. Add logging and rate limits for each.
7. Mark untrusted input: any content written by others (emails, tickets, web pages, shared files) is data, never instructions. List each such source for the AI Red Teaming Plan.
## Output Format
```markdown
# Agent Spec
**Agent:** [name] | **Version:** [number] | **Spec owner:** [role]
## Goal and done-state
[Goal] | Done when: [final state checked in the system]
## Tools and permissions
| Tool | Why the goal needs it | Permission (read / write / scope) | Narrower option considered |
|---|---|---|---|
| [tool] | [reason] | [permission] | [option] |
## Stop conditions
| Condition | Limit | Then |
|---|---|---|
| Steps / time / spend per session / repeated failure / blocker | [set by: name, date] | Stop and hand to [role] |
## Approval points (from Human-in-the-Loop Design)
| Action | Trigger | Approver |
|---|---|---|
| [action] | [trigger] | [role] |
## Never alone
| Action | Logged where | Rate limit |
|---|---|---|
| [action] | [log] | [team sets] |
**Untrusted input (data, never instructions; listed for the red team):** [emails, tickets, web pages, shared files]
## Decision
[Named person] sets the session budget and signs the never-alone list by [date].
```
## Done When
- The done-state is a checkable system state, and every tool has a reason and the narrowest permission that works
- Every stop condition ends in a handoff, and every limit is set by a named person or left blank
- Every untrusted input source is listed for the red team
## Quality Bar
- A prompt instruction is never the only control on a never-alone action: name the permission, approval or filter that enforces it
- No budget, limit or cost figure is filled by Claude; [set by: name, date] until a person sets it
- Agents that act on people (messages, accounts) go through approval points; no agent action profiles users
- Claude drafts the limits; a named person sets the budget and signs the never-alone list
## Next
Run aipm-tool-descriptions (Tool Descriptions) to write the tools the spec allows.
## About the makers
This pack is made by Polar Bear, a consultancy built by ex-McKinsey founders with a dream to make AI work for People, not instead of them. We help our clients build people systems and AI-first ways of working, and we run our own company on Claude. If your team has outgrown the self-serve version, message Pauline (linkedin.com/in/paulinebertry).