Adversarially pressure-test a product, startup, or feature idea to find whether it is a real business or an interesting idea. Runs a seven-test battery — problem severity, willingness to pay, switching cost, founder-market fit, timing, distribution, and the cheapest falsifying experiment — grading every claim on an evidence ladder that separates what people say from what they pay for. Use when someone asks to validate, stress-test, pressure-test, sanity-check, gut-check, poke holes in, red-te...
Scanned 9/6/2026
Install to Claude Code
npx -y skills add tmoody1973/pressure-test --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of pressure-test?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/tmoody1973-pressure-test)More formats (shields.io, HTML) on the badges page.
---
name: pressure-test
description: Adversarially pressure-test a product, startup, or feature idea to find whether it is a real business or an interesting idea. Runs a seven-test battery — problem severity, willingness to pay, switching cost, founder-market fit, timing, distribution, and the cheapest falsifying experiment — grading every claim on an evidence ladder that separates what people say from what they pay for. Use when someone asks to validate, stress-test, pressure-test, sanity-check, gut-check, poke holes in, red-team, or "be honest about" an idea, a startup, a pivot, a feature bet, or a business model; when they ask "would anyone pay for this", "is this a real business", "why won't this work", or "what am I missing"; or before committing engineering time to a new product direction. Deliberately falsifying — for expansive market research, competitive landscaping, and positioning, use startup-validator instead.
---
# Pressure Test
A battery for finding out whether an idea is a business. It is built to
disconfirm, not to encourage. If it cannot find the reason this fails, that is
information; if it can, that is cheaper to learn now.
**Announce it:** "Using pressure-test to run the battery on [idea]."
---
## The one idea this skill is built on
> **What people say is not evidence. What people do is evidence. What people pay
> is the best evidence you can get before you have customers.**
Every question below is a way of dragging a claim down from talk to behaviour.
The battery's job is to locate each claim on the ladder, then say plainly how
much weight it can hold.
---
## The Evidence Ladder
Grade every load-bearing claim. Cite the rung by number in your output. This is
the spine of the whole skill — if you do nothing else, do this.
| Rung | What it is | What it is worth |
|---|---|---|
| **0** | Founder conviction, logic, analogy, "obviously people want this" | Nothing. Not evidence. |
| **1** | Compliments. "I love this." "Cool idea." | Nothing — actively misleading, because it feels like progress |
| **2** | Stated intent. "I'd definitely buy that." "We'd pay $50/mo." | Weak. Systematically inflated — see the bias discount below |
| **3** | Adjacent spend. They already pay for something in this job | Moderate. Proves a budget exists |
| **4** | Costly signal. Time booked, a colleague introduced, data shared, a pilot scheduled, reputation staked | Strong. They spent something real |
| **5** | Money committed. Deposit, pre-order, signed LOI, paid pilot | Very strong |
| **6** | Repeat money. They paid, used it, and paid again | Definitive |
**The bias discount.** Stated willingness to pay reliably overstates real
payment. Meta-analyses put the inflation between roughly **1.35x and 3x**,
worse for yes/no questions, worse for anything socially flattering to endorse,
worse for public goods. Never quote a rung-2 number without applying and naming
a discount. See `references/research.md` for the studies and their limits.
**Rung 3 is the quiet workhorse.** "What do you spend on this today?" beats
"would you pay?" every time, because it is archaeology rather than prophecy.
---
## The Battery
Seven tests. **T2, T4 and T5 are the original three** this skill was built
around; the rest close gaps that kill more products than those three combined.
| | Test | The question it answers | Load-bearing? |
|---|---|---|---|
| **T1** | **Problem Severity** | Is this a fire, or a preference? | Yes — gates everything |
| **T2** | **Willingness to Pay** | Will they pay, or just approve? | Yes |
| **T3** | **Switching Cost** | Is it enough better to be worth leaving what they use? | Yes |
| **T4** | **Founder–Market Fit** | Is this edge real, or is it familiarity? | No — changes odds, not viability |
| **T5** | **Timing** | Why now, and what if it is not now? | Yes |
| **T6** | **Distribution** | How do the first hundred find you, twice? | Yes |
| **T7** | **The Kill Shot** | What is the cheapest thing that could prove this wrong this week? | Output, not a gate |
**They chain — they do not average.** T1 gates T2: if the problem is not real,
willingness-to-pay evidence is noise about a fantasy. T2 gates T6: you cannot
compute whether a channel is affordable until you know what a customer is worth.
**Never average a fatal flaw into a decent overall score.** One FAIL on a
load-bearing test caps the whole verdict, whatever else passed.
Full operational detail — the questions to ask, what counts as a pass, the
traps in each — is in `references/battery.md`. **Read it before running the
battery.** This file is the doctrine; that one is the instrument.
---
## Modes
| Invocation | What runs |
|---|---|
| `/pressure-test <idea>` | Full battery, T1–T7 |
| `/pressure-test quick <idea>` | T1, T2, T6 — the three that kill fastest — plus T7 |
| `/pressure-test wtp <idea>` (or `timing`, `founder`, `switching`, `distribution`) | One test, in depth |
| `/pressure-test interview <idea>` | Skip the battery; produce a customer-conversation script from `references/interviewing.md` |
Works on a live product, a pivot, or a feature bet — not only a napkin idea.
For a shipped product, prefer rung 5–6 evidence you already have and say so.
### Before you start
If the idea is one vague line, ask **at most three** clarifying questions, and
only ones whose answer would change a verdict. Good ones: *who exactly feels
this*, *what do they do about it today*, *who signs the cheque*. Then proceed on
stated assumptions rather than interrogating further. Do not stall the battery
waiting for perfect input — run it on what you have and mark the gaps `UNKNOWN`.
---
## Verdicts
### Per test
`PASS` · `WEAK` · `FAIL` · `UNKNOWN`
`UNKNOWN` is a real, respectable answer and must be used when evidence is
absent. Do not launder a guess into a `WEAK`.
### Overall — pick exactly one
| Verdict | What it means |
|---|---|
| **REAL** | Load-bearing tests carry rung 4+ evidence. Go. |
| **PROMISING, UNPROVEN** | Logic holds; evidence sits at rung 2–3. The next experiment is named and cheap. This is where most honest early ideas land. |
| **INTERESTING IDEA, NOT A BUSINESS** | Something people like and nobody buys. Usually a T1 or T2 failure. |
| **BROKEN AS SPECIFIED** | A specific load-bearing test fails in a way iteration on this shape will not fix. Name the test and the fix that would be required. |
| **CANNOT ASSESS** | Not enough information, and inventing it would be worse than saying so. List exactly what is missing. |
Never hedge between two. Pick one and defend it.
---
## Rules
**On evidence**
1. Enthusiasm is not validation. Compliments are rung 1 — treat them as noise.
2. Every load-bearing claim gets a rung number, stated inline.
3. Never quote stated WTP without discounting it and saying you did.
4. Weak evidence must be called weak, in those words. Do not dress rung 2 as rung 4.
5. **"No competitors" is bad news until proven otherwise.** An empty market usually
means no market, not open field. The incumbent is nearly always the status quo —
a spreadsheet, an intern, an email thread, or doing nothing — and it is winning.
**On honesty about what you do not know**
6. **Never invent a market number.** Tag every figure: `[sourced: …]`,
`[estimate: derivation shown]`, or `[unknown]`. A fabricated TAM is worse than
an admitted gap, because it gets repeated in a pitch and then to an investor
who checks.
7. If you have not searched and the number matters, say "I have not verified
this" rather than producing a confident range.
8. Bottom-up beats top-down. `customers × price × frequency` you can defend beats
a 1% slice of a market you cannot.
**On stance**
9. Do not open with a compliment. Start with the finding.
10. Before any verdict, write **the strongest case against** the idea — the
version a smart sceptic who wants you to succeed would make.
11. Challenge every claimed advantage. Ask whether it is durable, or whether a
funded competitor buys it in a quarter.
12. If someone else is better positioned to build this, say who and why.
13. Do not confuse passion with positioning, familiarity with insight, or a
growing market with a reachable one.
14. **Every verdict ships with its falsifier**: the specific observation that
would flip it. A verdict nothing could change is an opinion.
**On what not to do**
15. Do not average away a fatal flaw.
16. Do not soften a FAIL into a WEAK because the person seems invested.
17. Do not recommend "more customer discovery" as the next step without
specifying who, how many, which question, and what result would kill it.
18. Do not produce a score out of 100. False precision invites gaming and hides
which test actually failed.
---
## Output shape
```
PRESSURE TEST — <idea, in one line as you understood it>
Mode: full | quick | single Assumptions I ran on: <…>
## Strongest case against
<the sceptic's best shot, 3–6 sentences, made properly — not a straw man>
## The battery
T1 Problem severity ......... PASS/WEAK/FAIL/UNKNOWN — <one line> [rung N]
T2 Willingness to pay ....... …
T3 Switching cost ........... …
T4 Founder–market fit ....... …
T5 Timing ................... …
T6 Distribution ............. …
<a short paragraph per test: the finding, the evidence and its rung, the trap>
## Verdict
**<one of the five>**
<why, in three or four sentences, naming which test drove it>
## What would change my mind
<the specific observation that flips the verdict>
## The kill shot (T7)
The riskiest assumption: <…>
The cheapest test: <who, how many, what you ask, what you measure>
Cost: <time and money> Timebox: <days>
Kill condition: <the result at which you stop>
```
Adapt the shape to the medium — a Slack answer does not need headings — but
never drop the rung citations, the strongest case against, the falsifier, or the
kill shot. Those four are the skill.
---
## Reference files
| File | Read it when |
|---|---|
| `references/battery.md` | **Always, before running.** The seven tests in operational detail |
| `references/interviewing.md` | Designing customer conversations, or the user's evidence is all rung 1–2 |
| `references/research.md` | Citing a finding, or challenged on where a number comes from |
---
## The failure mode of this skill
An assistant that agrees. Sycophancy is the default failure of a model asked to
evaluate something its user is emotionally invested in, and it is the exact
failure this skill exists to prevent. If your draft output would make the
founder feel good and would not change what they do on Monday, you have written
a compliment with structure. Start again.
The kindest thing this skill can do is find the flaw while it is still cheap.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!