Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Hypothesis Design

ASecurity

Builds an answer-first plan by stating disprovable hypotheses and, for each, the "what would prove me wrong" test and the exact analysis that settles it, so the team tests instead of boiling the ocean.

8 stars
0 votes
0 copies
1 views
Added 9/19/2026
ai-agentsgotesting

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add andreworia/claude-consulting-skills --skill hypothesis-design --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Hypothesis Design?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Hypothesis Design
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/andreworia-hypothesis-design/badge)](https://www.skillsdirectory.com/skills/andreworia-hypothesis-design)

More formats (shields.io, HTML) on the badges page.

Download with Pro
Files
SKILL.md
---
name: hypothesis-design
description: Builds an answer-first plan by stating disprovable hypotheses and, for each, the "what would prove me wrong" test and the exact analysis that settles it, so the team tests instead of boiling the ocean.
---

# Hypothesis Design

## When to use
Use this when the team is about to gather data with no point of view, at risk of analyzing everything and concluding nothing. It is the right skill when you want an answer-first approach: state the likely answer up front as a set of disprovable hypotheses, then spend effort only on the analyses that could confirm or kill them. It is the counter to boiling the ocean.

## What it does
It turns a question into a small set of sharp, disprovable hypotheses and pairs each with its decisive test. A good hypothesis is a specific claim that could be wrong, and for which you can name the evidence that would disprove it. The skill produces a prioritized hypothesis set and, for each, the single analysis that would settle it, so the workplan is lean and pointed.

## Method
The skill runs hypothesis-led problem solving, the practice of leading with a provisional answer and testing to disprove it.

1. State the provisional answer. Given the key question and what the team already knows, write the best current guess at the answer in one sentence. This is deliberately committal. An answer-first plan needs an answer to test.

2. Decompose the answer into component hypotheses. Break the provisional answer into the two to five claims that must each hold for the answer to be right. Draw these from the issue tree branches where one exists. Each becomes a hypothesis.

3. Write each hypothesis to be disprovable. A well-formed hypothesis is:
   - Specific: it names the driver, the direction, and where it applies.
   - Falsifiable: you can state a result that would make it false.
   - Consequential: if true, it changes the recommendation.
   Rewrite vague hypotheses ("marketing could improve") into sharp ones ("the binding growth constraint is activation in the mid-market segment, not top-of-funnel demand").

4. Define the disproof condition. For each hypothesis, complete the sentence "I would abandon this hypothesis if I saw ___." Naming the disproof up front is what prevents the analysis from quietly turning into a hunt for supporting evidence.

5. Name the decisive analysis. For each hypothesis, identify the single analysis or piece of evidence that most cheaply distinguishes true from false. Prefer the test that could kill the hypothesis fastest. Note the data required and roughly how hard it is to get.

6. Prioritize the hypothesis set. Rank hypotheses by two factors: how central each is to the answer, and how uncertain it currently is. Test the central, uncertain ones first. If the lead hypothesis is disproved early, revise the provisional answer and re-derive, rather than pressing on.

7. Set the branch logic. State what you will conclude and do next under each outcome (hypothesis holds, hypothesis fails). This makes the plan a decision tree, not a data-collection list.

## Inputs
- The key question (from problem definition) and any issue tree.
- What the team already believes the answer might be.
- A sense of which claims are most uncertain.

## Output format
A hypothesis plan:
- Provisional answer: the one-sentence committal guess.
- Hypothesis set: each hypothesis as a specific, falsifiable claim.
- For each hypothesis: the disproof condition ("I would abandon this if..."), the single decisive analysis, and the data it needs.
- Priority order: the hypotheses ranked by centrality and uncertainty.
- Branch logic: what happens to the recommendation under each outcome.

## Example
Key question: why is revenue growth stalling, and what recovers it fastest?

Provisional answer: growth is stalling because the largest segment is churning faster than new logos replace it, so the fastest recovery is retention, not acquisition.

Component hypotheses:
- H1: net revenue retention in the top segment has fallen below replacement. Disproof: retention is flat or rising in that segment. Decisive analysis: cohort retention by segment over eight quarters. Data: billing records by cohort.
- H2: churn is driven by a specific unmet need, not price. Disproof: churned accounts cite price as the primary reason. Decisive analysis: structured exit interviews coded by reason. Data: churned-account contacts.
- H3: acquisition is healthy, so it is not the constraint. Disproof: new-logo volume is also falling. Decisive analysis: new-logo trend by source.

Priority: H1 first (central and uncertain). Branch logic: if H1 holds, the recommendation centers on retention and H2 tells us the lever; if H1 fails, the provisional answer is wrong and the plan pivots to acquisition, re-deriving from H3.

Attribution

andreworiaandreworia
View sourceMore from andreworia →
SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

693161 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →