Improve how well an agent can verify its own work in this repo: setup and worktree flows, debug access, test speed, end-to-end QA loops, screenshots, log access. Use when the user says "improve agent dx", "what do you need to verify your work", "make this repo agent-friendly", "verification loop", "you keep saying it should work", or runs /theo-mode dx.
Scanned 9/19/2026
npx -y skills add al3rez/theo-mode --skill agent-dx --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Agent Dx?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/al3rez-agent-dx)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: agent-dx
description: >
Improve how well an agent can verify its own work in this repo: setup and
worktree flows, debug access, test speed, end-to-end QA loops, screenshots,
log access. Use when the user says "improve agent dx", "what do you need to
verify your work", "make this repo agent-friendly", "verification loop",
"you keep saying it should work", or runs /theo-mode dx.
---
# Agent DX
Goal: after this session, an agent in this repo can go from "made a change" to
"proved the change works" without a human. The deliverable is tooling and
docs, committed.
## Step 1: answer the question honestly
Before touching anything, write down in `NOTES.md` the answer to:
> "For the kind of change most PRs here make, what would I need to run to be
> sure it works, and can I run it right now?"
Check the last 20 merged PRs (`gh pr list --state merged --limit 20`) to see
what layer they touch: UI, API, internal functions, config. That layer is
where the verification loop matters.
Then try to actually verify one recent PR's change. Every place you get
stuck is an item for step 2. Common ones:
- No way to launch the app headless. No way to screenshot it.
- Tests take 8 minutes so nobody runs them; or they need a database that
needs a secret that needs a person.
- Logs go to a service the agent cannot read.
- The dev server and the test runner fight over the same port.
- A `.env.example` that is missing half the keys.
- Seed data does not exist, so every flow starts with an empty screen.
- No worktree-safe setup: two checkouts share a `node_modules` or a database.
## Step 2: build the missing pieces
Fix the top items, cheapest first. Typical outputs:
- `scripts/dev-up.sh` (or a `make dev`) that gets from clean clone to running
app with seed data in one command, idempotent, works in a git worktree.
- A `/run-<app>` skill with a driver: `chromium-cli` script for web, `curl`
smoke script for a server, tmux wrapper for a TUI. Screenshots to a known
path. (The `run-skill-generator` skill builds this if available.)
- A fast test target: `test:changed` or `test:unit` that runs in under 30s.
- Debug access: a `DEBUG=` flag documented in CLAUDE.md, a `logs/` tail
command, a `--json` output mode on the CLI.
- A `CLAUDE.md` section titled "How to verify a change" with the exact
commands. Keep it under 20 lines.
Each piece must be run by you, in this session, before it is documented.
## Step 3: prove it
Re-do the verification from step 1 using only the new tooling. Time it.
Report: what it took before (or "impossible"), what it takes now, what is
still missing and why.
## Cost guardrail
Do not build a full e2e suite. Build the one path that covers what PRs
actually change, and make it fast. Breadth is a follow-up the user can ask
for.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!