Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

stackfit

ASecurity

Recommends third-party services, APIs and SDKs for the specific repository in front of you rather than in the abstract. Reads the codebase's frameworks, data layer, deployment target and existing vendors, measures SDK health and finds real implementations on GitHub, scores candidates on a weighted rubric, and reports blast radius, effort, risks and an integration plan. Use this whenever someone asks which service, provider, SDK or API to use for a feature - auth, payments, subscriptions, bill...

2 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentstypescriptpythongobashsqlnextjsgitapidatabase

Works with

cliapi

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add angellane/stackfit --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of stackfit?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for stackfit
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/angellane-stackfit/badge)](https://www.skillsdirectory.com/skills/angellane-stackfit)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: stackfit
description: Recommends third-party services, APIs and SDKs for the specific repository in front of you rather than in the abstract. Reads the codebase's frameworks, data layer, deployment target and existing vendors, measures SDK health and finds real implementations on GitHub, scores candidates on a weighted rubric, and reports blast radius, effort, risks and an integration plan. Use this whenever someone asks which service, provider, SDK or API to use for a feature - auth, payments, subscriptions, billing, file storage, email, push notifications, search, analytics, realtime, queues, vector or AI infrastructure - or asks to compare providers ("Stripe vs Paddle", "is Clerk worth it"), asks "what should I use for X", weighs build versus buy, asks what switching providers would cost, or wants an Architecture Decision Record for a technology choice. Use it even when the request sounds casual, like "how should I add billing to this app".
license: MIT
---

# StackFit

Developers rarely need to know the best payments API in general. They need to know
what to do in the repository open in front of them, given what it already uses,
where it deploys, and how much time they have. Generic advice is free everywhere;
the value here is entirely in the specificity.

So the order is fixed: **read the repo, then form an opinion.** An opinion formed
first and justified afterwards is a guess, and a scoring table wrapped around a
guess just makes it look rigorous.

## Three failure modes to design against

1. **Generic recommendation** - would read identically for a different codebase.
   Fix: every claim traces to a path, a dependency, or a measured fact.
2. **False precision** - a score of 87.3, a price recalled from training data.
   Fix: verify volatile facts or label them unverified.
3. **Ignoring hidden work** - the SDK call is ten lines; the webhook endpoint,
   idempotency, reconciliation job and migration are the actual project.

## Workflow

Script paths below are relative to this skill's directory; the repository being
analysed is the argument. The working directory is normally the developer's
project, so use the skill's own path for the script and `.` for the target.
Substitute the real path for `$SKILL_DIR`.

Commands are written as `python3`. If that fails - on Windows it is often a Store
stub that prints nothing or opens a prompt - retry with `python`, then `py -3`,
and use whichever works for the rest of the session. Write scratch files such as
`candidates.json` to the system temp directory (`$TMPDIR`, `%TEMP%`), never `/tmp`
on Windows and never into the developer's repository. If a script still won't run,
say which one and why in the report; don't silently skip its output.

### 1. Match depth to the question

- **Quick call** ("which email API?", "is Clerk overkill?") - analyse the stack,
  answer in a few paragraphs with one runner-up. No scoring table.
- **Full evaluation** ("help me pick a payments provider") - the whole workflow.
- **Migration** ("Firebase to Supabase, how bad?") - read
  `references/migration-analysis.md` instead; the effort profile is different.

Ask a question only when the answer changes the ranking. "Seat-based or usage-based
billing?" changes the winner; "what's your timeline?" usually doesn't and can be
inferred from the repo. Never open with a questionnaire.

### 2. Read the repository

```bash
python3 "$SKILL_DIR/scripts/analyze_stack.py" . --format text
```

Deterministic, offline, never reads real `.env` files. In a monorepo, target the
specific package - a root scan mixes stacks and produces a muddled picture.

Then **confirm by opening actual files.** The analyzer sees that Prisma is
installed; it cannot see that it's used in one file and raw SQL everywhere else, or
that the "auth" in the manifest is a half-finished prototype. Two well-chosen files
change recommendations more than any amount of dependency listing.

Four findings eliminate candidates outright:

- **Runtime model.** Serverless and edge hosts cannot run persistent workers or
  hold websockets. Options assuming a long-lived process are disqualified, not
  merely penalised.
- **Existing vendors.** A platform already in the stack starts ahead - one fewer
  vendor, bill, SDK and auth model.
- **Data layer.** Decides how painful new tables are and whether a provider's
  schema expectations fit.
- **Webhook and job capability.** Missing infrastructure is usually the largest
  cost line in the whole integration.

### 3. Frame the decision

Name the two or three things that actually decide this case - usually the runtime
constraint, the existing vendor, and one product requirement (global payouts,
HIPAA, self-hosting, EU residency). A hard requirement is a filter, not a criterion:
options failing it get excluded with a reason, not scored 2/5.

### 4. Shortlist

Three to five candidates. More means the research wasn't narrowed, and a long
undifferentiated list hands the decision back to the developer.

Always weigh two options that get forgotten:

- **The vendor already in the stack.** Before recommending a specialist for auth,
  storage, search or realtime, check whether the platform already present handles
  it adequately. Often it does. When the specialist is genuinely better, say what
  specifically is better - "more powerful" is not a reason to add a vendor.
- **Building it.** Build when the problem is bounded and the operational surface is
  small: presigned S3 uploads, Postgres full-text search, a notifications table,
  flags in a config table. Buy when it hides a state machine or a compliance
  surface: billing, auth, deliverability, payments. The test isn't difficulty, it's
  how many edge cases carry real consequences when handled wrong.

Read `references/domains/<domain>.md` for that domain's landscape, deciding factors
and hidden requirements. Playbooks exist for `auth`, `payments`, `storage`, `email`,
`sms`, `notifications`, `search`, `analytics`, `realtime`, `ai`, `jobs`, `database`,
`cms`, `feature-flags` and `observability`.

**This skill is not limited to those.** The workflow applies to any feature a
developer wants to add - e-signature, maps, video, i18n, PDF generation, anything.
When no playbook exists, derive the same four things yourself before shortlisting:
what the realistic candidates are, which one or two factors actually decide it, what
the domain quietly drags in (new infrastructure, a schema change, a compliance
surface, an ongoing sync obligation), and which facts are volatile enough to need
verifying. Say plainly that you worked without a playbook, so the developer knows
the landscape came from reasoning rather than a curated list.

### 5. Measure what can be measured

```bash
python3 "$SKILL_DIR/scripts/github_probe.py" probe stripe @paddle/paddle-js
python3 "$SKILL_DIR/scripts/github_probe.py" find "nextjs stripe subscription" --language TypeScript
```

`probe` turns `sdk_quality` and `ecosystem_maturity` into measurements: whether the
package is deprecated, when it was last published, whether the repo is archived, how
recently it was pushed. A deprecated or archived SDK is decisive and easy to miss
from memory. `find` surfaces real implementations worth reading for integration
patterns - examples, not endorsements, since popular starter repos are often
outdated or built for a different stack.

The probe reads package registries first and treats GitHub as enrichment, because
the GitHub API allows 60 requests/hour unauthenticated and shared IPs exhaust that
fast. Set `GITHUB_TOKEN`, or have `gh` logged in, for 5000/hour. When the quota is
gone it returns registry data and says so - carry that limitation into the report
rather than presenting partial data as complete.

Stars measure popularity, not maintenance or fit. An official vendor SDK with 400
stars is often better maintained than a community wrapper with 4000.

### 6. Verify anything that moves

Architecture facts are stable (Stripe is webhook-driven; Postgres has full-text
search). Pricing, free-tier limits, rate limits, SDK versions and regional
availability are volatile.

For volatile facts: fetch the vendor page and cite it, or label the claim -
*"roughly 2.9% + 30c per training data, verify before committing."* An unlabelled
stale price is worse than none, because it gets budgeted against. If web access
isn't available, say so once and label everything volatile.

### 7. Score

Build a candidates file (`score_candidates.py --template`) using the anchored
definitions in `references/scoring-rubric.md` - anchors are what make two runs
agree. All ten criteria are oriented so 5 is favourable, so cost, effort and
lock-in are scored as `cost_efficiency`, `implementation_speed`, `portability`.

```bash
python3 "$SKILL_DIR/scripts/score_candidates.py" "$TMPDIR/candidates.json" --profile default --format markdown
```

Profiles: `default`, `ship-fast`, `cost-sensitive`, `enterprise`, `long-haul`. Use
what the developer signalled; otherwise `default`, and name the profile, since the
weighting is an assumption they can override.

Three checks keep the answer honest. Carry their results into the report rather
than quietly dropping them:

- **Margin** under five points is a tie. Say so and give a concrete tiebreaker.
  Always report the margin as the number the scorer printed ("12.4 points"),
  never as an adjective like "clear". If you describe the top two as
  interchangeable or near-identical, that is a tie whatever the number says -
  call it one. If scoring didn't run, write "not scored" and why, rather than
  asserting a margin you didn't measure.
- **Sensitivity** - if the winner changes under other profiles, state the
  condition: "Stripe if flexibility matters more, Paddle if you don't want to
  handle sales tax."
- **Evidence audit** - scores with no reason, facts never verified. Fix, don't ship.

If the scoring contradicts your instinct, work out which is wrong. Never adjust
scores so a preferred option wins; change the recommendation instead, or admit the
criteria don't capture what matters here.

### 8. Work out what it touches

```bash
python3 "$SKILL_DIR/scripts/impact_scan.py" . --domain payments   # --list-domains
python3 "$SKILL_DIR/scripts/impact_scan.py" . --domain "e-signature for contracts"
```

The domain argument takes free text and resolves it, so a feature description works
as well as a keyword. Anything unrecognised falls back to a generic integration
profile rather than failing - it still finds credential handling, existing API
clients, config and data models, which is most of what any integration touches.

Reports touchpoint files by role, infrastructure readiness, and an effort band with
its inputs shown. Treat the band as a floor for someone who already knows the
codebase. Turn it into specifics with `references/impact-analysis.md`: schema,
dependencies, env vars, ordered phases. "Roughly twelve files" is weak;
"`prisma/schema.prisma` needs a Subscription model and
`app/api/webhooks/stripe/route.ts` doesn't exist yet" is what makes it credible.

## Report structure

Full evaluations use `assets/report-template.md`. Sections, in order: recommendation
and headline numbers; detected stack; options table with margin and sensitivity; why
the winner over the runner-up; what it touches; integration plan; risks including
lock-in; what would change this recommendation.

That last section is the most useful one a year later - it tells the team when to
revisit rather than re-litigating from scratch. Drop any section that would be
padding for the question asked.

## Output economy

Developers read this output; every wasted line costs them attention and costs the
recommendation clarity. Aim for density, not brevity for its own sake.

- **Lead with the decision.** No preamble, no restating the question. The first
  line is the answer.
- **Don't narrate the process.** Show findings, not "I ran the analyzer, then
  searched GitHub". Tool output is evidence, not content to summarise.
- **Never state a fact twice.** If it's in the stack table, don't repeat it in prose.
- **Tables for comparison, prose for reasoning.** Prose listing parallel attributes
  should have been a table.
- **Cut any sentence true of any repo.** Generic vendor praise, boilerplate caveats
  and restatements of what the developer just told you are the main waffle sources.
- **No closing summary.** The recommendation was the first line; repeating it adds
  nothing.
- Rough budgets: quick call under 250 words, full report under 900. Exceeding them
  is fine when the content is dense - but check that it is.

## Judgment rules

**Ground every claim** in a path, dependency, measurement, fetched source, or a
labelled assumption.

**Recommend one thing.** Present the ranking, commit to an answer. Declare genuine
ties, then give the tiebreaker.

**Respect what works.** Suggesting a rewrite of functional code because another
provider scores marginally higher is bad advice; migration cost lands on people who
did nothing wrong.

**Name the lock-in.** What leaving costs: which data is portable, which code is
provider-shaped, roughly how long an exit takes. Developers can accept lock-in
knowingly; they shouldn't discover it later.

**Numbers carry their basis.** Effort bands say who they assume. Prices carry a
source or a label. Scores carry evidence.

## Anti-patterns

Recommending before reading the repo. Eight options with no ranking. Pricing from
memory stated as current. An architecture the deployment target can't run - a
persistent worker on Vercel, a websocket server on static hosting. Ignoring a vendor
already in the manifest. Scoring everything 4/5 so nothing separates. Treating the
weighted score as the decision rather than as shown reasoning.

## Follow-ups

Offer at most one, when it fits: an **ADR** (`references/adr-template.md`) after a
decision; **migration analysis** for "should we move from X to Y"; or starting
**phase one** of the integration plan.

## Files

| Path | Read when |
|---|---|
| `references/scoring-rubric.md` | Before scoring - anchored 0-5 definitions |
| `references/domains/<domain>.md` | While shortlisting, for that domain only |
| `references/impact-analysis.md` | Turning scan output into a plan |
| `references/migration-analysis.md` | Provider-to-provider migrations |
| `references/adr-template.md` | Writing an ADR |
| `assets/report-template.md` | Writing a full report |

| Script | Purpose |
|---|---|
| `analyze_stack.py` | Stack, vendors, deployment, capabilities, constraints |
| `github_probe.py` | SDK health from registries and GitHub; find real integrations |
| `score_candidates.py` | Weighted scoring, margin, sensitivity, evidence audit |
| `impact_scan.py` | Touchpoints, infrastructure readiness, effort band |

All are standard-library Python 3.8+ and read-only. The first, third and fourth are
offline and deterministic; `github_probe.py` is the only one making network calls,
and degrades to partial results rather than failing. Run `--help` on any of them. If
a script fails, read the repo directly and say so - the logic matters more than the
tooling.

Attribution

angellaneangellane
View sourceMore from angellane →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →