Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Eve Forge

BSecurity

Forge a personalized eve agent for a business end-to-end — absorb the business's artifacts, author the eve agent/ dir, validate, deploy to Vercel (or a VPS), smoke-test against ground truth, register, and evolve. The deterministic core is three safety gates distilled from a driven benchmark (BRO-1677): deploy-safety (never ship auth:none), validate (eve info --json → 0 diagnostics + tools registered), and smoke (drive the deployed agent, assert vs a ground-truth example). USE WHEN building/on...

4 stars
0 votes
0 copies
0 views
Added 9/27/2026
ai-agentspythongobashnode

Works with

cli

Security Analysis

B75/100
criticalExfiltrates credentials via HTTP — exact pattern from Snyk ToxicSkills study
criticalSends environment variables or credentials to an external URL

Pro scans all 16 files and shows the line behind each finding

Scanned 9/27/2026

$npx -y skills add broomva/skills --skill eve-forge --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Eve Forge?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Eve Forge
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/broomva-eve-forge/badge)](https://www.skillsdirectory.com/skills/broomva-eve-forge)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: eve-forge
category: orchestration
version: 1.0.2
author: broomva
description: >-
  Forge a personalized eve agent for a business end-to-end — absorb the
  business's artifacts, author the eve agent/ dir, validate, deploy to
  Vercel (or a VPS), smoke-test against ground truth, register, and evolve.
  The deterministic core is three safety gates distilled from a driven
  benchmark (BRO-1677): deploy-safety (never ship auth:none), validate
  (eve info --json → 0 diagnostics + tools registered), and smoke
  (drive the deployed agent, assert vs a ground-truth example). USE WHEN
  building/onboarding an eve agent for a tenant, "forge an eve agent",
  "deploy an eve agent", "onboard <business> onto eve", or when the
  Claude-Code orchestrator/forge must turn absorption inputs into a running
  eve operator. NOT FOR benchmarking frameworks (that's a one-off), running
  the operator itself (the forge builds it; the operator runs cheap on eve),
  or non-eve agent frameworks.
tags:
  - eve
  - agent-forge
  - meta-agent
  - vercel
  - deployment
  - claude-agent-sdk
  - dogfood
  - agent-substrate
compounding: each forged tenant is a versioned agent/ dir + tenant-spec.json (authored-agents-as-data); re-onboard/evolve = a diff
provenance:
  - docs/reports/2026-07-04-life-vs-eve-benchmark.html
  - research/entities/concept/eve-agent-orchestrator.md
  - BRO-1677
---

# eve-forge — turn a business into a deployed eve agent

The orchestrator/forge (a Claude Agent SDK program) produces a **deployed,
tenant-scoped eve agent** from a business's absorption inputs. This skill
encodes the real eve workflow + every trap learned dogfooding it, and gates
the consequential steps so the forge cannot repeat the benchmark's mistakes.

**Latent vs deterministic split:**
- **Latent (agent judgment, this SKILL.md):** `absorb` (Word template + transcript
  + examples → `tenant-spec.json`) and `author` (write the eve `agent/` files in
  the business's voice from the templates).
- **Deterministic (`scripts/`, tested):** `preflight` (Node ≥ 24), `deploy-safety`
  (never ship unlocked auth), `validate` (eve info clean), `smoke` (output vs
  ground truth). Precision work lives in code; the latent space invokes it.

## The 8-stage pipeline

| # | Stage | How | Gate |
|---|---|---|---|
| 1 | **Absorb** | read the template + 1–2 transcripts + 2–3 filled examples → write `tenant-spec.json` (see `references/templates/tenant-spec.example.json`) + a ground-truth `truth.json` (required substrings + case-scoped `forbidden`). **Stage these OUTSIDE the tenant dir** — `eve init` refuses a non-empty target ("has no package.json") | latent |
| 2 | **Scaffold** | `python3 scripts/eve_forge.py preflight` **then** `nvm use 24 && npx eve@latest init <slug>`, then move `tenant-spec.json`/`truth.json` in | **preflight blocks if Node < 24** (the npx trap) |
| 3 | **Author** | fill from `tenant-spec.json`: copy `references/templates/{fill_document,send_document}.ts` → `agent/tools/`; write `agent/instructions.md` (business voice + **"strip HTML comments from the output"**); **EDIT the scaffolded `agent/channels/eve.ts`** — remove `placeholderAuth()` → `auth: [vercelOidc(), localDev()]` (never `none()`). Do NOT hand-write `defineChannel`; the scaffold already ships `eveChannel` | latent |
| 4 | **Validate** | `npx eve info --json \| python3 scripts/validate.py --expect-tools fill_document,send_document` (validate.py strips eve's banner + reads the real dict-`diagnostics`/`status` schema) + `npm run typecheck` | **0 diagnostic errors + tools registered, or iterate** |
| 5 | **Deploy** | `python3 scripts/eve_forge.py gate agent/` (point at the **`agent/` dir**, not the project root) **before** `vercel deploy --scope <team>`; use the **production alias** (the raw URL 302s to SSO) | **deploy-safety denies if auth not locked** |
| 6 | **Smoke** | drive the deployed agent (see **§Smoke against a locked channel**) → `python3 scripts/smoke.py --output <filled.txt> --truth truth.json` | **assert vs ground truth (evidence-gated)** |
| 7 | **Register** | commit `freelance/<slug>/` (agent dir + `tenant-spec.json` + `truth.json` + `smoke-receipt.json`: `{url, verdict, coverage, at}`); report URL + evidence | tenant = versioned data |
| 8 | **Evolve** | owner draft→approve corrections → forge proposes a diff to `instructions.md`/skills | **propose → test vs fixtures → owner-approve → commit** |

## The deploy-safety gate (the incident-derived check)

A benchmark run shipped a **public, `auth: none()`, Gateway-billed** eve endpoint —
anyone could spend credits. This skill makes that structurally unreachable. In the
Claude-Code orchestrator, wire it as a **PreToolUse hook** that runs
`scripts/deploy_safety.py <agent_dir>` before any `vercel deploy` and **denies the
tool call on a non-zero exit**. Rule (prod): the channel `auth:` array must contain a
real authenticator (`vercelOidc`) and must NOT contain `none()`/`placeholderAuth()`;
a lone `localDev()` is dev-only. Fail-closed if no `auth:` array is found.

## Smoke against a locked channel

A correctly-locked channel returns **401** to anonymous callers — so the smoke driver must
authenticate (the only reason a naive smoke "worked" in the benchmark was the `auth: none()`
incident). On the Vercel deploy:

```bash
vercel env pull /tmp/<slug>.env --environment=production --scope <team>   # mints VERCEL_OIDC_TOKEN
TOKEN=$(grep VERCEL_OIDC_TOKEN /tmp/<slug>.env | cut -d= -f2- | tr -d '"')
# POST a turn — payload field is `message` (NOT `input`); returns 202 + sessionId:
curl -s -XPOST "https://<slug>.vercel.app/eve/v1/session" -H "Authorization: Bearer $TOKEN" \
     -H 'content-type: application/json' -d '{"message":"<transcript>"}'
# GET the stream (it long-polls at session.waiting → cap it), extract the fill_document output:
curl -s --max-time 60 "https://<slug>.vercel.app/eve/v1/session/<sessionId>/stream?startIndex=0" \
     -H "Authorization: Bearer $TOKEN" > /tmp/<slug>.stream
```
Delete `/tmp/<slug>.env` after (it holds a live token).

## Deterministic scripts

- `scripts/eve_forge.py preflight` — Node ≥ 24 or fail (the `npx eve init` trap).
- `scripts/eve_forge.py gate <agent_dir> [--info info.json --expect-tools a,b]` — runs deploy-safety (+ validate) as one pre-deploy gate.
- `scripts/deploy_safety.py <agent_dir> [--env prod|dev]` / `--stdin` — the auth-lock check.
- `scripts/validate.py --expect-tools a,b` (reads `eve info --json` on stdin) — diagnostics + tools.
- `scripts/smoke.py --output <file> --truth <json>` — deployed-output vs ground-truth (strips HTML comments; `truth.json` `forbidden` encodes case-scoped negatives, e.g. no `"bloodwork"` for a non-senior patient).

## Gotchas (see `references/gotchas.md`)

Node-24 hard requirement (npx silently uses the wrong Node) · non-TTY `eve dev` errors
(scaffold succeeds, the auto-dev-launch fails) · fail-closed default auth + Vercel
Deployment-Protection SSO (use the production alias, keep auth locked) · eve not
auto-detected as a Vercel framework ("No framework detected" → its native Agent-Runs
observability may not activate) · AI-Gateway auth works via project OIDC at runtime
(zero keys). Pin the eve + claude CLI versions (eve is beta).

## Anti-rationalization

| Excuse | Reality |
|---|---|
| "Deploy it, I'll lock auth later." | The deploy-safety gate is binary and PreToolUse-wired. `auth: none()` never reaches prod. |
| "eve info was clean when I wrote it." | Re-run `validate` after every author edit; typecheck drift is silent. |
| "It looks right, ship it." | Smoke-test against the business's own ground-truth example, or it's prose, not evidence. |
| "npx eve init failed weirdly." | Run `preflight` — it's the Node-24 trap 9 times out of 10. |

## References

- `references/gotchas.md` — the full benchmark gotcha list.
- `references/templates/` — the 4 eve agent file templates the author stage fills.
- `research/entities/concept/eve-agent-orchestrator.md` — the orchestrator design + provenance.
- `docs/reports/2026-07-04-life-vs-eve-benchmark.html` — the driven benchmark this distills.

Attribution

broomvabroomva
View sourceSee grades on GitHubMore from broomva →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →