
Claude Skills by The-Utopia-Studio
github.com/The-Utopia-StudioAudits infrastructure setup — database config, caching, performance, and scaling patterns. Provides concrete detection commands and improvement steps. Use when the user asks about "infrastructure", "database setup", "caching", "scaling", "performance", or "architecture". Don't use for CI/CD setup (use deployment-engineer), monitoring (use monitoring-setup), or security (use security-auditor).
Create technical and product diagrams — architecture, flowchart, sequence, state machine, ER / data model, timeline, swimlane, quadrant, nested, tree, layer stack, venn, pyramid — as standalone HTML files with inline SVG. Ships with a neutral editorial skin and a first-run gate that prompts users to customize the style guide (colors, fonts) from their own website before generating. Includes annotation-callout primitive and optional sketchy variant.
Strip designs to their essence by removing unnecessary complexity. Great design is simple, powerful, and clean. Use when the user asks to simplify, declutter, reduce noise, remove elements, or make a UI cleaner and more focused.
Design graphic assets with Efecto — presentations, pitch decks, event posters, email headers, blog images, open graph cards, business cards, resumes, menus, infographics, invitations, newsletters, and documents. Use when asked to "design a poster", "create a pitch deck", "make a presentation", "design a business card", or any graphic design task. Requires Efecto MCP server.
Design web pages and app UIs with Efecto — create sessions, build layouts with JSX and Tailwind CSS, manage artboards, and push real-time design changes via MCP tools. Use when asked to "design a page", "build a landing page", "create a website", "design a dashboard", "make a UI", or any visual design task. Requires Efecto MCP server.
Turn a validated wedge into the scope you can SCORE — a one-sentence job, 20 pass/fail golden cases drawn from real artefacts, an L0–L4 autonomy level with a failure taxonomy and a derived acceptable failure rate per mode, and a cost-per-outcome budget to the cent. Fires on "spec the build", "define scope", "scope the v1", "write the spec", "how do we know it works". Not for the component pipeline or effort split (use compound-system-architecture), not for pilot price / terms / commercial suc...
Records visual proof while testing UI behavior — screen recording with structured test/assertion annotations — then posts the video and a results summary to the PR and tracker issue. Use whenever a UI change needs verifiable evidence that it works, instead of prose claims.
Weights a pile of "validation" by what people DID, not what they said. Takes interview quotes, sign-ups, LOIs, pilots, payments and places each on one ladder — money moved 1.0, behaviour observed 0.7, artefact shown 0.5, verbal commitment 0.3, opinion 0.1 — then returns a weighted evidence table, the weight of the load-bearing claim, and for every low signal the cheapest probe that raises it a rung. Fires on "how strong is this signal", "score the interview", "did they actually validate", "we...
Split a body of expertise into the tell-able procedures (explicit) and the show-only judgment (tacit), and flag the tacit half as the defensible product. Fires when a fellow says "what's teachable vs judgment", "codify the expertise", "split explicit from tacit", "which part of our know-how is defensible", "what can we write down vs what's in their head", or "turn our expertise into a product spec". Returns an Explicit/Tacit Ledger: every piece of know-how classified by one test — could a str...
Assess where ONE fellow sits on the Icarus ladder — Literate → Practitioner → Operator → Frontier → Author — by checking which EXIT ARTEFACT each level requires actually EXISTS and can be cited, never by how ready the fellow feels. Fires on "how's this fellow doing", "assess a fellow", "what level am I / is this fellow at", "am I ready to level up", "level up this fellow". Output is a filled level assessment: the ladder table with a cited artefact and evidence rung per rung, the assigned leve...
Routes a fellow through the Icarus module by situation type, not by tearing down one idea. Fires on "where do I start", "I already have a product / traction — which stages apply to me", "onboard me to Icarus", "which stages should I skip", "I have a mature product, what's my path". Classifies the fellow as Type A (blank page), B (traction, no moat), or C (mature product) on the evidence ladder, then returns a keep / trim / subtract / leap stage ledger and a think:build:test ratio tied to that...
Analyzes profitability per customer, product, or transaction to determine business model viability and scalability. Covers CAC, LTV, contribution margin, cohort analysis, and growth-readiness assessment. Use when evaluating business model viability, validating startup metrics (CAC, LTV, payback period), making pricing decisions, comparing business models, or when user mentions unit economics, CAC/LTV ratio, contribution margin, customer profitability, or break-even analysis.
Turns a concept into the cheapest concrete artefact a human can react to, via the no-code make-sequence: Crazy 8s → paper/Miro flow → digital mock → clickable hybrid in v0 or Figma Make, built in one afternoon. Fires on "mock it up", "run Crazy 8s", "clickable prototype", "turn this idea into something I can click", "get from a rough idea to a prototype in an afternoon". Output is a Clickable-Prototype Plan: the ONE thing the mock must provoke a reaction to, the fidelity band, the no-code cei...
Scores a specific concept across the four build-risk lenses — Desirability, Usability, Feasibility, Viability — where each lens is graded ONLY by its named tool (onion+JTBD+Kano for Desirability, a watched usability observation for Usability, a dev spike for Feasibility, ICE anchored to measured value for Viability) and run by its dual-track owner (Product/Design discovery beside Engineering delivery). Fires on "is this desirable / usable / feasible / viable", "should we build it", "run the f...
Iteratively improves a PR (GitHub) or MR (GitLab) until Greptile gives it a 5/5 confidence score with zero unresolved comments. Triggers Greptile review, fixes all actionable comments, pushes, re-triggers review, and repeats. Use when the user wants to fully optimize a PR/MR against Greptile's code review standards.
Fires when a fellow needs to decide how an AI product is stopped from doing the wrong thing — "design the guardrails", "when does a human sign off", "how do we handle failures / bad outputs", "what confidence threshold should we auto-approve at", "where do we put the human in the loop". Returns a guardrail spec: every failure mode placed on a cost-of-error × volume matrix, a three-layer stack (rules in code → confidence threshold → human sign-off) sized per mode, a derived confidence threshol...
Hugging Face Hub CLI (`hf`) for downloading, uploading, and managing models, datasets, spaces, buckets, repos, papers, jobs, and more on the Hugging Face Hub. Use when: handling authentication; managing local cache; managing Hugging Face Buckets; running or scheduling jobs on Hugging Face infrastructure; managing Hugging Face repos; discussions and pull requests; browsing models, datasets and spaces; reading, searching, or browsing academic papers; managing collections; querying datasets; con...
Create distinctive, production-grade frontend interfaces with high design quality. Generates creative, polished code that avoids generic AI aesthetics. Use when the user asks to build web components, pages, artifacts, posters, or applications, or when any design skill requires project context. Call with 'craft' for shape-then-build, 'teach' for design context setup, or 'extract' to pull reusable components and tokens into the design system.
Ink terminal renderer for json-render that turns JSON specs into interactive terminal UIs. Use when working with @json-render/ink, building terminal UIs from JSON, creating terminal component catalogs, or rendering AI-generated specs in the terminal.
Detects and connects development tools — Slack with GitHub, Linear with Git, error tracking with notifications. Provides step-by-step setup procedures for each integration. Use when the user asks to "connect tools", "link Slack", "set up notifications", "integrate Linear", or "connect my tools". Don't use for monitoring setup (use monitoring-setup), deployment (use deployment-engineer), or security (use security-auditor).
Invent an original product concept before the machine hands you the generic one — onion the stated idea to its invariant core need, diverge and keep the strange child, anchor every concept to your proprietary YODA corpus, then run the generic-prompt test (if the one-line prompt a competitor would type reaches your concept, it isn't yours yet). Returns an invented concept + rationale. Fires on "what should we actually build", "invent the solution", "make it non-obvious", "give me a concept a c...
Reduces a job to its three irreducible currencies — information moved, decisions made, liability transferred — after deleting every tool, vendor, product, and org-chart role name from the description. Fires on "what job is really being done here", "strip this down to the primitive", "what's the primitive job", "take the tool names out and tell me the underlying job", or when a fellow describes a workflow thick with product and team names and wants the tool-independent job beneath it. Outputs ...
Fires when a fellow asks which numbers matter for a launched product — "what metrics should we track", "what should our North Star be", "is our retention any good / does the curve flatten", "our MAU is up-and-to-the-right, is that real or vanity", "what's our cost per outcome". Returns a metric scorecard: one customer-centric North Star, an AARRR one-metric skeleton, the flattening-retention truth test that GATES the North Star, and cost-per-outcome priced to the cent. Do NOT fire to design t...
Sequence a validated wedge into a defensible moat as four DATED gates (wedge → usage → lock-in → data advantage), forced through an incumbent-veto sentence ("X won't copy this because ___") and three monthly falsifiers; returns a filled moat canvas. Fires on "what's the moat", "how is this defensible", "will this compound", "design the moat", "will competitors just copy this / how do we stay defensible as they show up". Not for whether one wedge gets adopted now (use wedge-five-questions — ru...
Configures error tracking (Sentry), product analytics (PostHog), health monitoring, and alerting. Provides verification commands and concrete output templates. Use when the user asks to "set up monitoring", "add error tracking", "configure Sentry", "add analytics", or "set up alerts". Don't use for deployment (use deployment-engineer), security (use security-auditor), or integration linking (use integration-linker).
Falsification pass for the load-bearing beliefs under a chosen plan or wedge. Restates each assumption as a null hypothesis (the belief is FALSE), designs the single cheapest observation whose failure would disprove it, and ranks every assumption by P(wrong) x impact-if-wrong so the fellow shoots at the most-likely- fatal belief first. Fires on "what has to be true", "how would I disprove this", "what's the riskiest assumption", "what could kill this", "which belief do I test first". Output i...
Creates concise, decision-ready product specifications (one-pagers and PRDs) that align stakeholders on problem, solution, users, success metrics, and constraints. Use when proposing new features/products, documenting product requirements, creating concise specs for stakeholder alignment, pitching initiatives, scoping projects before detailed design, capturing user stories and success metrics, or when user mentions one-pager, PRD, product spec, feature proposal, product requirements, or brief.
Pushes interfaces past conventional limits with technically ambitious implementations — shaders, spring physics, scroll-driven reveals, 60fps animations. Use when the user wants to wow, impress, go all-out, or make something that feels extraordinary.
Fires when a fellow wants to test a workflow on paper before building it — "paper test", "sketch it", "paper prototype", "sketch probe", "draw the flow and check they can follow it". Output is a hand-drawn workflow sketch plus a structured read-out naming what was legible and where the decision lived. NOT for choosing which probe to run (that is `probe-matrix`), and NOT for testing demand, trust, or willingness to pay — a sketch lies about all three; route those to `wizard-of-oz-probe` or `co...
Sizes the per-unit prize of a piece of work from first principles by comparing what it is priced at today against its theoretical floor once AI does the automatable part. Fires on "is this a big enough problem", "how big is the prize", "size the opportunity from first principles", "what's the physics floor", "is the gap big enough", or when a fellow has a unit of work and its current per-unit cost and wants a build/walk verdict. Outputs a filled floor/gap calc sheet: token-cost line + judgmen...
Turn a would-be pilot into a six-term paid-pilot sheet — scope, price (paid or prepaid), the signed data-rights clause, success metrics, kill criteria, and conversion terms priced now — and refuse to call anything a pilot unless all six terms are non-empty. Fires on "pilot terms", "structure the deal", "paid pilot", "pilot term sheet", "structure the pilot so it isn't a free trial". Not for brainstorming or validating which revenue model to bet on (that is `monetization-strategy`, exploratory...
Draws the ownership line for an AI build. For every component it decides one of three things: Daedalus (the studio platform) builds it once and the fellow inherits it, the fellow builds it because it is their moat, or it is a rented commodity. Fires on "build vs buy", "build or use the platform", "what does Daedalus give me", "do we build our own eval harness / RAG / router", "what's ours vs the studio's", "should we build our own model". Returns a build/buy boundary map that splits each comp...
Routes the ONE question a fellow needs answered to the cheapest probe that is HONEST about that question, using the probe honesty contract (paper/sketch, Wizard-of-Oz, concierge, agent-concierge — each honest about some things and lying about others). Fires on "how do I test this cheaply", "which experiment", "what's the cheapest way to learn X", "which probe", "how do I validate this". Output is a probe selection + plan: the question, the chosen probe, why it is honest about this question, w...
Numeric go/no-go gate that scores ONE product problem 1-5 on eight evidence-backed dimensions (frequency, budgeted pain, severity, data exhaust, structural persistence, buyer clarity, wedge sharpness, founder asymmetry), sums to /40, and returns pass (>=32) / redesign (28-31) / kill (<28). Fires on "should I build this", "score this problem", "go or no-go", "is this problem good enough", "rate this problem". Output is a filled 8-row scorecard with a money-or-behaviour citation on every row. N...
Restates a product idea as ONE decision a named human makes, then quantifies how that decision is compressed — before→after on a single axis of time, effort, or autonomy (e.g. six minutes → thirty seconds). Fires on "what's the product here", "frame the problem", "what decision are we changing", "state this as a decision", "what are we actually changing for the user". Output is a filled Compressed-Decision Statement: the decision (never a feature), its one owner, the before→after compression ...
Walks one validated problem up a load-bearing stack — validated problem → vision → strategy → product vision → North Star → OKRs → dual-track roadmap — and BLOCKS any layer from resting on an unvalidated problem below it. Fires on "frame the business", "vision to roadmap", "what's the strategy", "turn this validated problem into a roadmap", "give me the vision, North Star and OKRs". Output is a filled frame stack where every layer inherits the problem's evidence rung and every roadmap item la...
PM skill for Claude Code, Codex, Cursor, and Windsurf. Diagnoses SaaS metrics, critiques PRDs, plans roadmaps, runs discovery, coaches PM career transitions, pressure-tests AI product decisions, and designs PLG growth strategies. Seven knowledge domains, 12 templates, 40+ frameworks, and an opinionated interaction style that labels assumptions and names tradeoffs.
Makes a repository AI-native by creating AGENTS.md, .cursorrules, and documentation that helps any AI agent understand the codebase. Follows the WHAT/WHY/HOW framework from best practices. Use when the user asks to "make repo AI-ready", "add cursor rules", "create AGENTS.md", "set up for AI coding", or "make agents work better". Don't use for deployment (use deployment-engineer), monitoring (use monitoring-setup), or code organization (use repo-structurer).
Guides validation of ideas before full development using pretotyping (fake doors, concierge MVPs, Wizard of Oz) and prototyping at appropriate fidelity (paper, clickable, coded) to test assumptions about demand, pricing, and feasibility. Use when testing ideas cheaply before building, choosing prototype fidelity, running experiments to validate assumptions, or when user mentions prototype, MVP, fake door test, concierge, Wizard of Oz, landing page test, smoke test, or asks "how can we validat...
Tones down visually aggressive or overstimulating designs, reducing intensity while preserving quality. Use when the user mentions too bold, too loud, overwhelming, aggressive, garish, or wants a calmer, more refined aesthetic.
This skill should be used when the user asks about Central Station threads, community discussions, support questions, feature requests, or wants to search Railway's community knowledge base. Use for queries like "search central station", "find threads about", "what are people asking about", "recent support threads", or "central station topics".
This skill should be used when the user wants to add a database (Postgres, Redis, MySQL, MongoDB), says "add postgres", "add redis", "add database", "connect to database", or "wire up the database". For other templates (Ghost, Strapi, n8n, etc.), use the templates skill.
This skill should be used when the user wants to push code to Railway, says "railway up", "deploy", "deploy to railway", "ship", or "push". For initial setup or creating services, use new skill. For Docker images, use environment skill.
This skill should be used when the user wants to manage Railway deployments, view logs, or debug issues. Covers deployment lifecycle (remove, stop, redeploy, restart), deployment visibility (list, status, history), and troubleshooting (logs, errors, failures, crashes, why deploy failed). NOT for deleting services - use environment skill with isDeleted for that.
This skill should be used when the user wants to add a domain, generate a railway domain, check current domains, get the URL for a service, or remove a domain.
This skill should be used when the user asks "what's the config", "show me the configuration", "what variables are set", "environment config", "service config", "railway config", or wants to add/set/delete variables, change build/deploy settings, scale replicas, connect repos, or delete services.
This skill should be used when the user asks about resource usage, CPU, memory, network, disk, or service performance. Covers questions like "how much memory is my service using" or "is my service slow".
This skill should be used when the user says "setup", "deploy to railway", "initialize", "create project", "create service", or wants to deploy from GitHub. Handles initial setup AND adding services to existing projects. For databases, use the database skill instead.
This skill should be used when the user wants to list all projects, switch projects, rename a project, enable/disable PR deploys, make a project public/private, or modify project settings.
This skill should be used when the user asks about Railway features, how Railway works, or shares a docs.railway.com URL. Fetches up-to-date Railway docs to answer accurately.