Skip to content
Back to skills

Forge Builder

ASecurity

Universal 7-phase metric pipeline builder. Autonomously discovers upstream data, reasons about metric requirements as a data analyst, engineers pipeline code as a data engineer, and deploys as a data architect. Handles GitHub issue intake with spec-challenge quality gates, source table discovery, metric design, pipeline generation (Daily + Cumulative), schema DDL, ALTER scripts, unit tests, deployment, PR creation, and issue lifecycle management. Triggers on "metric builder", "build metric", ...

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 7, 2026
ai-agentspythongoshellsqlnodeexpressrailstestinggitdocumentation

Works with

  • cli

Security analysis

A100/100

Pro scans all 19 files and shows the line behind each finding

Scanned October 7, 2026

npx -y skills add tewarishishir/dataforge --skill forge-builder --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Forge Builder?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Forge Builder
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tewarishishir-forge-builder/badge)](https://www.skillsdirectory.com/skills/tewarishishir-forge-builder)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: forge-builder
description: >-
  Universal 7-phase metric pipeline builder. Autonomously discovers upstream
  data, reasons about metric requirements as a data analyst, engineers pipeline
  code as a data engineer, and deploys as a data architect. Handles GitHub issue
  intake with spec-challenge quality gates, source table discovery, metric design,
  pipeline generation (Daily + Cumulative), schema DDL, ALTER scripts,
  unit tests, deployment, PR creation,
  and issue lifecycle management. Triggers on "metric builder", "build metric",
  "build metrics", "new metric", "add metric", "metric from issue",
  "forge-builder", "dataforge/builder", "end to end metric", "metric pipeline",
  "explore this table for metrics", "what metrics can I build from",
  "discover metrics from table", "create metric from source".
metadata:
  status: "beta"
  agent-tools: [read, search, edit, runCommands, githubRepo]
  human-description: Builds a metric pipeline end to end from a GitHub issue or plain-language request, through schema, code, tests, and deployment.
---

# forge-builder

Universal, end-to-end metric pipeline builder. Takes a GitHub issue, natural language request, or metric definitions reference and produces every artifact needed: pipelines, schemas, ALTER scripts, tests, deployment, PR, and issue lifecycle management.

Before starting, read the required context:

- `dataforge.manifest.resolved.yaml` -- all project-specific values (catalogs, domains, zones, naming, grains, downstream integrations, orchestration, runner, spark_conf)
- [forge-lint](../forge-lint/SKILL.md) -- column naming, SQL formatting, Python conventions, quality checklist

If the resolved manifest does not exist, or if the seed `../../assets/dataforge.manifest.yaml` has a newer modification time, invoke [forge-init](../forge-init/SKILL.md) to regenerate it before proceeding. If the seed `../../assets/dataforge.manifest.yaml` contains legacy full-manifest keys (e.g., `catalogs`, `domains`), read it directly as a pre-resolved manifest.

## Manifest Contract

This skill reads `../../assets/dataforge.manifest.resolved.yaml` for every business-specific value. The resolved manifest is auto-generated by [forge-init](../forge-init/SKILL.md) from the seed `../../assets/dataforge.manifest.yaml` via Unity Catalog CLI discovery and codebase scanning. It supplies: project, platform, dbt, catalogs, naming, grains, domains, downstream, orchestration, runner, spark_conf, entity_structure, and tool_name_overrides. No business constants are hardcoded in this skill.

Throughout this document, manifest references use dot-path notation (e.g., `manifest.project.repository`). The agent reads the resolved YAML file and resolves these paths at runtime.

## Agent Toolchain

| Skill | When to Invoke |
| :--- | :--- |
| [forge-lint](../forge-lint/SKILL.md) | Every file modification (mandatory standards) |
| [forge-dbx](../forge-dbx/SKILL.md) | Authentication, data discovery queries, schema change deployment, pipeline runs |
| `gh` CLI | Issue intake, spec challenges, progress comments, pull request creation, issue closure |

## Skill Resolution

Before starting Phase 1, verify that every skill in the Agent Toolchain is present. If a skill file is missing from the expected path, fetch it from the shared registry.

### Check for installed skills

Look for each skill's `SKILL.md` at its expected relative path (e.g., `../forge-lint/SKILL.md`). If the file does not exist, the skill is not installed.

### When to fetch

| Situation | Action |
| :--- | :--- |
| Required skill missing at expected path | Fetch by name, then re-read |
| Agent needs a capability not covered by current skills | Search registry, fetch if found |
| Skill exists but may be outdated | Check for updates |

Do NOT proceed past Phase 1 if `forge-lint` or `forge-dbx` are missing -- these are blocking dependencies.

## Reference Files (load on-demand per phase)

| File | Phases | Content |
| :--- | :--- | :--- |
| [system-memory-store.md](references/system-memory-store.md) | All | v2 JSON schema, field reference, checkpoint table, resume protocol |
| [gate-discovery.md](references/gate-discovery.md) | 1-3 | Verification checklists for intake, discovery, and design phases |
| [gate-engineering.md](references/gate-engineering.md) | 4-5 | Verification checklists for schema and pipeline phases; cross-phase column alignment |
| [gate-deployment.md](references/gate-deployment.md) | 6-7 | Verification checklists for testing, deployment, and delivery |
| [gate-recovery.md](references/gate-recovery.md) | On failure | Self-healing patterns, runtime error catalog, edge cases, data quality red flags |
| [phase-1-intake.md](references/phase-1-intake.md) | 1 | Intake mode procedures (A/B/C/D), Universal Hypothesis Boundary, Scoping Decision Tree |
| [phase-1-issue-intake.md](references/phase-1-issue-intake.md) | 1, 7 | `gh` invocation patterns, spec challenge, comment templates, task list creation |
| [phase-2-discovery.md](references/phase-2-discovery.md) | 2 | Mandatory 5-query sequence, query failure halt protocol, Hypothesis Resolution Log, Discovery Presentation Gate |
| [phase-3-design.md](references/phase-3-design.md) | 3 | SQL expression design, join strategy, PoC query procedure, user approval gate |
| [phase-4-schema.md](references/phase-4-schema.md) | 4 | Schema DDL generation procedures, file naming, entity folder structure, ALTER strategy, cross-phase column contract |
| [phase-5-pipeline.md](references/phase-5-pipeline.md) | 5 | Pipeline code generation procedures, file generation order, dry-run validation, job definition wiring, multi-source integration |
| [phase-6-testing.md](references/phase-6-testing.md) | 6 | Test SQL patterns, lint quality gate execution, _update_count exclusion rules |
| [phase-7-deployment-playbook.md](references/phase-7-deployment-playbook.md) | 7 | Branch naming, PR body, rollback, table version history, zone summary, Schema Sign-Off gate |
| [phase-7-delivery.md](references/phase-7-delivery.md) | 7 | README updates, development history table, delivery summary, monitoring baseline |
| [code-templates/spark_python.md](references/code-templates/spark_python.md) | 4, 5 | Daily/Cumulative Python, Schema SQL, ALTER script, job definition templates |
| [code-templates/dbt.md](references/code-templates/dbt.md) | 4, 5 | dbt model, sources.yml, schema.yml templates |
| [on-demand-removal.md](references/on-demand-removal.md) | On demand | Deprecation protocol for retired metrics |
| [on-demand-batch.md](references/on-demand-batch.md) | On demand | Multi-tool batch build rules |

## Workflow Router

| User Intent | Entry Point | Notes |
| :--- | :--- | :--- |
| "build metrics for X", "new metric", "add metric" | Phase 1 (Intake Mode B) | Natural language intake |
| "build metrics from issue 1234", "metric from issue" | Phase 1 (Intake Mode A) | GitHub issue intake |
| "build metrics from definitions SQL for X" | Phase 1 (Intake Mode C) | Definitions SQL reference |
| "explore this table for metrics", "what metrics can I build from X" | Phase 1 (Intake Mode D) | Discovery-driven intake via forge-dbx |
| "resume build", "continue issue 1234", "what's left" | Session Management -> Resume Protocol | Reads `.build-state/{build_id}.json` |
| Orphan build detected (pipeline code exists, no build-state file) | Session Management -> Orphan Detection | Inspect files, infer phase, write build-state, resume |
| "deprecate X", "retire metric X" | [on-demand-removal.md](references/on-demand-removal.md) | Deprecation protocol |
| "create PR for this build" | Phase 7 (Git and PR) | Uses `gh pr create --body-file` per deployment patterns |

If the user's intent is ambiguous, ask up to 3 targeted questions before routing.

---

## 1. Session Management

This skill spans 7 phases across 14+ files. Use three sessions to avoid context exhaustion. Mandatory checkpoints exist after every phase (1-7), so any session boundary is a safe hand-off point.

| Session | Phases | Deliverables | Context Budget |
| :--- | :--- | :--- | :--- |
| Discovery and Design | 1-3 | Scoping document, data discovery report, metric design spec | ~30% |
| Engineering | 4-5 | Schema DDL, pipelines, job definitions | ~40% |
| Testing and Deployment | 6-7 | Tests, quality gate, PR, deployment, delivery summary | ~30% |

### Context Budget

Quality degrades predictably as context fills:

| Context Usage | Quality | Action |
| :--- | :--- | :--- |
| 0-30% | Peak | Thorough, comprehensive -- ideal for discovery and design |
| 30-50% | Good | Solid engineering work, confident decisions |
| 50-70% | Degrading | Shortcuts begin -- if mid-phase, complete and checkpoint |
| 70%+ | Poor | STOP. Write memory store, hand off to next session |

**Spawn subagents when:**

- Data discovery requires querying 4+ source tables (parallel discovery agents)
- Phase 7 backfill generation is complex (multi-table entity with 3+ CTEs)
- Phase 5 pipeline code touches 5+ files simultaneously
- Any phase would push past 50% context usage

At each session boundary, summarize all artifacts and build state so the next session can resume without re-reading the full pipeline.

### Build Memory Store

Every build persists its state to `skills/forge-builder/.build-state/{build_id}.json`. See [system-memory-store.md](references/system-memory-store.md) for the v2 schema template, `build_id` format, checkpoint field table, write protocol, and resume rules.

**Checkpoint timing is non-negotiable.** Write the checkpoint immediately after completing each phase — before beginning any work on the next phase. Do not defer checkpoints to the end of a build session.

### Phase Entry Protocol (Hook)

At the start of every phase N > 1, execute this shell command:

    python skills/forge-builder/verify.py phase-entry {build_id} {N}

- **Exit code 0:** Script prints `Phase N entry check: PASS`. Proceed with phase work.
- **Exit code 1:** Script prints the halt reason. Do not proceed. Follow the printed instruction.

This command is non-negotiable. Every phase below has a **Hook** instruction that invokes it. If the script is not executed, the phase did not start correctly. The script reads the build-state file and asserts `completed_phases` contains `N-1` — this verification is deterministic and cannot be self-attested.

**If verify.py is missing or cannot import:** The script is absent from the expected path or the Python environment cannot resolve it. Do not halt the build on this condition alone. Fall back to manual checkpoint assertion: read `.build-state/{build_id}.json` directly and confirm `completed_phases` contains `N-1`. If the field is present and N-1 is listed, proceed. If not, treat as Case B below. Append to `notes`: `"[PROTOCOL] verify.py unavailable at phase-entry {N}. Manual checkpoint assertion used."` Continue with the phase work.

**If the script halts:** Determine the cause:

| Case | Condition | Action |
| :--- | :--- | :--- |
| **A — Missed checkpoint write** | Phase N-1 ran this session; its output data exists in context; only the file write was skipped | Write the Phase N-1 checkpoint retroactively. Append to `notes`: `"[PROTOCOL] Phase N-1 checkpoint missed. Written retroactively at {ISO timestamp}."` Run the verify.py command again. |
| **B — Phase N-1 never ran** | No Phase N-1 output exists in context | Do NOT write a fabricated checkpoint. Re-run Phase N-1 from its entry point. |

If there is any doubt — if no Phase N-1 outputs are present in the current session context — treat it as Case B.

Each phase below has a one-line **Hook** instruction that invokes this protocol. Do not skip it.

### Resume Protocol

At the start of every session, before Phase 1:

1. Construct the `build_id` from the user's request.
2. Check for `skills/forge-builder/.build-state/{build_id}.json`.
3. If found, read the file and **verify content for every phase listed in `completed_phases`** — search pipeline code files in `artifacts` for metric aliases (e.g., `rg "metric_name" path`). Downgrade any claimed-complete phase with no evidence in actual file content.
4. Present the verified summary: Tool, Issue, Verified Completed vs. Claimed but Unverified, Next phase, Metrics, Blocking.
5. Apply idempotency and drift detection rules from [system-memory-store.md](references/system-memory-store.md). Continue the pipeline.
6. If the file does not exist, start fresh from Phase 1.

### Orphan Detection

An orphan build is a partially completed pipeline where artifacts exist on disk but no `.build-state/` file is present.

1. Scan the entity's folder under its domain directory (per `manifest.domains`) for existing files: Daily Python, Cumulative Python, Schema SQL, ALTER SQL.
2. For each artifact found, **verify content — not just file existence.** Grep for the specific metric aliases. A file exists but lacks the metric = phase incomplete.
   - Schema DDL contains the metric column → Phase 4 completed
   - Daily/Cumulative `run_query()` contains the metric alias → Phase 5 completed
   - Downstream integration files contain the metric name → Phase 6 completed
3. Write a new `.build-state/` file with `completed_phases` set to the **content-verified** phases.
4. Present the inferred state to the user for confirmation before resuming via the standard Resume Protocol.

> For multi-tool batch builds, see [on-demand-batch.md](references/on-demand-batch.md). For metric removal and deprecation, see [on-demand-removal.md](references/on-demand-removal.md).

---

## Phase 1: Intake and Scoping

Supports four intake modes. A GitHub issue is the primary expected path.

### Phase 1 Entry Check

**This check runs before any external call, issue fetch, or intake mode execution.** Its purpose is to detect an existing build for this `build_id` before overwriting it or duplicating start state.

1. Construct the candidate `build_id`:
   - Mode A: `unknown--issue-{number}` (tool_name not yet resolved) or `{tool_name}--issue-{number}` if tool_name is explicitly known from the user's request.
   - Modes B/C/D: `{tool_name}--{YYYYMMDD-HHmmss}` (timestamp from user request context).
2. Check if `.build-state/{build_id}.json` exists.
3. **File found and `completed_phases` includes `1`:** This is an active resume. Run the Resume Protocol from Session Management. Do not re-execute Phase 1.
4. **File found, `completed_phases` is empty, `current_phase: 0`:** The build crashed during Phase 1 after the start state write but before the Phase 1 checkpoint. Do NOT overwrite the existing file. Continue Phase 1 execution from where context was lost — skip the Pre-Phase 1 Start State Write (the file already exists).
5. **File not found:** Proceed to Pre-Phase 1 Start State Write.

### Pre-Phase 1: Start State Write

**Runs only when the Phase 1 Entry Check confirms no existing build-state file.** This is a safety anchor: if Phase 1 fails mid-execution (e.g., a `gh` call times out after a comment is posted), the resume protocol finds this file and does not restart from scratch.

- For Mode A: construct `build_id` as `{tool_name}--issue-{number}` using the issue number from the user's request. If `tool_name` is not yet known, use `unknown` as a placeholder — it will be resolved and the file will be renamed at the Phase 1 checkpoint (see tool-name resolution below).
- For Modes B/C/D: construct `build_id` as `{tool_name}--{YYYYMMDD-HHmmss}`. Use the current UTC timestamp — do not estimate; look it up from system context.
- Write the file with `status: "intake_started"`, `current_phase: 0`, `completed_phases: []`, `blocking: null`.
- Acquire the lock file per the write protocol in [system-memory-store.md](references/system-memory-store.md) before writing.

**Tool-name resolution (Mode A only):** When `tool_name` is resolved from `unknown` during Phase 1 (after the issue is fetched and the tool is identified), execute this rename sequence as part of the Phase 1 checkpoint:

1. Write the full Phase 1 checkpoint to the correct `build_id` (with the resolved `tool_name`, e.g., `incident--issue-1726`).
2. Delete the `unknown--issue-{number}.json` placeholder file.
3. Overwrite the lock file to reflect the new `build_id`.

Do not leave a file named `unknown--issue-{number}.json` on disk after Phase 1 completes. If the Phase 1 checkpoint write fails before deletion, the stale `unknown` file is acceptable as a crash artifact — the Resume Protocol will find the correct file by `build_id` construction on the next session.

> **Read on entry:** Load [phase-1-intake.md](references/phase-1-intake.md) for all intake mode procedures (A/B/C/D), the Universal Hypothesis Boundary, and the Scoping Decision Tree. Load [phase-1-issue-intake.md](references/phase-1-issue-intake.md) when executing Mode A for `gh` invocation patterns, the Minimum Metric Spec checklist, and the Spec Challenge Comment template.

**Artifact:** Scoping Document (in-memory): tool_name, metric_name(s), entity table(s), business definitions, grain requirements, temporal requirements, GitHub issue number (if Mode A), build path.

**Checkpoint 1 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 1 row. Verify: `verify.py schema-check {build_id}`.

---

## Phase 2: Data Discovery (Analyst Mode)

**Entry:** Phase 1 scoping document complete.

**Hook:** Execute `python skills/forge-builder/verify.py phase-entry {build_id} 2` via Shell. Nonzero exit = hard stop.

> **Read on entry:** Load [phase-2-discovery.md](references/phase-2-discovery.md) for the mandatory 5-query sequence, query failure halt protocol, source table lookup patterns, grain feasibility rules, domain placement, Discovery Presentation Gate, and Hypothesis Resolution Log format.

**BLOCKING — authenticate first:** Use `forge-dbx` to authenticate to the platform before executing any query. Codebase search does not satisfy any mandatory query.

Execute the mandatory 5-query sequence per [phase-2-discovery.md](references/phase-2-discovery.md). All 5 must execute — do not skip any. If any query fails, apply the halt protocol from the reference.

**BLOCKING — Discovery Presentation Gate:** Present the complete Data Discovery Report to the user. Wait for acknowledgment before proceeding to Phase 3.

**BLOCKING — Hypothesis Resolution Log:** Evaluate every Phase 1 hypothesis against live query results per [phase-2-discovery.md](references/phase-2-discovery.md). Every hypothesis must appear — unconfirmed hypotheses may not appear in Phase 3 SQL.

**Deliverable:** Confirmed source table path, full column inventory, grain feasibility matrix (ACTIVE / INACTIVE / SUSPECT per grain), domain assignment, all data quality flags.

**VERIFY:** Run [gate-discovery.md](references/gate-discovery.md) Phase 2 checklist. All BLOCKING items must pass before continuing.

**Checkpoint 2 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 2 row. Verify: `verify.py checkpoint-written {build_id} 2` then `verify.py hypothesis-count {build_id}`. Both must pass.

---

## Phase 3: Metric Design (Analyst Mode)

**Entry:** Phase 2 data discovery report complete.

**Hook:** Execute `python skills/forge-builder/verify.py phase-entry {build_id} 3` via Shell. Nonzero exit = hard stop. Then execute `python skills/forge-builder/verify.py hypothesis-count {build_id}` to confirm all hypotheses are resolved before design begins.

> **Read on entry:** Load [phase-3-design.md](references/phase-3-design.md) for SQL expression design patterns, join strategy selection, proof-of-concept query procedure, and the user approval gate.

**Mandatory obligations:**

1. Design a SQL expression for every metric in scope — verb mapped to driving date column, CASE WHEN structure, filter conditions from Phase 2 Query 5 enumeration results only
2. Determine join strategy using the lint join pattern decision tree (forge-lint Section 5)
3. Execute one proof-of-concept query per metric via `forge-dbx` — do not skip metrics after the first. Zero-count results must be diagnosed before proceeding
4. Present the complete metric design to the user for approval — autonomous mode is the only exception

**BLOCKING — User Gate:** Do not proceed to Phase 4 until the user explicitly approves the design or autonomous mode was invoked.

**Deliverable:** Metric Design Specification: SQL expressions, join strategy, CTE structure, grain matrix, assumptions. Updated in memory store `metrics` array.

**Checkpoint 3 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 3 row. Verify: `verify.py checkpoint-written {build_id} 3`.

---

## Phase 4: Schema Architecture (Engineer Mode)

**Entry:** Phase 3 metric design approved.

**Hook:** Execute `python skills/forge-builder/verify.py phase-entry {build_id} 4` via Shell. Nonzero exit = hard stop.

> **Read on entry:** Load [phase-4-schema.md](references/phase-4-schema.md) for schema DDL generation procedures, build path routing, entity folder structure, ALTER script strategy, and the cross-phase column contract. Then load [code-templates/spark_python.md](references/code-templates/spark_python.md) or [code-templates/dbt.md](references/code-templates/dbt.md) based on `manifest.platform.pipeline_framework` for DDL templates and SQL syntax.

**Deliverable:** CREATE TABLE DDL files for each active grain × cadence (spark_python) or `schema.yml` entries (dbt). ALTER scripts for existing entities. `ALTER_LOG.md` created or updated.

**VERIFY:** Run [gate-engineering.md](references/gate-engineering.md) Phase 4 checklist.

**Checkpoint 4 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 4 row. Verify: `verify.py checkpoint-written {build_id} 4`.

---

## Phase 5: Pipeline Code (Engineer Mode)

**Entry:** Phase 4 schema complete.

**Hook:** Execute `python skills/forge-builder/verify.py phase-entry {build_id} 5` via Shell. Nonzero exit = hard stop.

> **Read on entry:** Load [phase-5-pipeline.md](references/phase-5-pipeline.md) for pipeline code generation procedures, file generation order, dry-run validation protocol, job definition wiring, and multi-source integration patterns. Then load [code-templates/spark_python.md](references/code-templates/spark_python.md) or [code-templates/dbt.md](references/code-templates/dbt.md) if not already loaded from Phase 4 for the actual Python and SQL templates.

**Deliverable:** Daily and Cumulative pipeline files passing dry-run validation. Job definition task entries (new entities only). All generated files verified against schema DDL column order.

**VERIFY:** Run [gate-engineering.md](references/gate-engineering.md) Phase 5 checklist.

**Checkpoint 5 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 5 row. Verify: `verify.py checkpoint-written {build_id} 5` then `verify.py content-verify {build_id}`. Both must pass.

---

## Phase 6: Testing and Validation (Architect Mode)

**Entry:** Phase 5 complete.

**Hook:** Execute the following three checks in order via Shell. Any nonzero exit = hard stop.

1. `python skills/forge-builder/verify.py phase-entry {build_id} 6` — confirms Phase 5 checkpoint is written
2. `python skills/forge-builder/verify.py artifacts-exist {build_id}` — confirms all Phase 4-5 artifact files exist on disk
3. `python skills/forge-builder/verify.py content-verify {build_id}` — confirms every implemented metric alias exists in the required artifact types for the pipeline framework (spark_python: `.py` + `.sql`; dbt: `.sql` + `.yml`; hybrid: whichever artifact types are present); a passing `artifacts-exist` with a failing `content-verify` means a previous session wrote build state without writing code

> **Read on entry:** Load [phase-6-testing.md](references/phase-6-testing.md) for test SQL patterns, lint quality gate execution procedures, and `_update_count` exclusion rules. See [gate-deployment.md](references/gate-deployment.md) for the Phase 6 verification checklist.

**Deliverable:** Test SQL produced or updated for every implemented metric. All `_update_count` metrics excluded from test assertions. Full lint quality checklist passed.

**VERIFY:** Run [gate-deployment.md](references/gate-deployment.md) Phase 6 checklist and full lint quality gate (all 4 categories).

**Checkpoint 6 of 7:** Write per [system-memory-store.md](references/system-memory-store.md) Phase 6 row. Verify: `verify.py checkpoint-written {build_id} 6`.

---

## Phase 7: Deployment and Delivery (Architect Mode)

**Entry:** Phase 6 quality gate passed.

**Hook:** Execute `python skills/forge-builder/verify.py phase-entry {build_id} 7` via Shell. Nonzero exit = hard stop.

> **Read on entry:** [phase-7-deployment-playbook.md](references/phase-7-deployment-playbook.md) for branch naming conventions, PR body template, Schema Sign-Off gate (Guardrail S5), rollback procedures, and zone dispatch order. [phase-7-delivery.md](references/phase-7-delivery.md) for README update procedures, development history table format, delivery summary template, and the post-deployment monitoring baseline. [phase-1-issue-intake.md](references/phase-1-issue-intake.md) for the progress comment and issue lifecycle patterns (Mode A: close the issue; Modes B/C: open one).

**Deliverable:** Branch created from `main`. PR opened via `gh pr create --body-file`. Schema changes deployed across all zones in `manifest.platform.deployment_order` (sequential, stop on failure). Pipeline loads executed after schema success per zone. GitHub issue updated and closed.

**Documentation updates (always):** Store-level README — update key numbers. Entity development history table — add row with current date, contributor, issue number, summary.

**Documentation updates (new entity only):** Update visualization or diagram node labels.

**Final Delivery Summary:** Present what was built (files, metrics, table paths), where it lives (file paths, catalog paths, PR link, issue link), what to verify (deployment status, PR review), and any outstanding items (monitoring).

**Post-deployment monitoring baseline (48 hours):**

- Row counts deviate >20% from source volume estimate → investigate pipeline run
- Any metric 100% zero across all dates → check entity ID column, date predicate, filter expression
- Zero rows in one zone while others are populated → pipeline run failure in that zone

**VERIFY:** Run [gate-deployment.md](references/gate-deployment.md) Phase 7 checklist.

**Checkpoint 7 of 7 (Final):** Write per [system-memory-store.md](references/system-memory-store.md) Phase 7 row. Release the lock file. Verify: `verify.py checkpoint-written {build_id} 7`.

---

## Data-Default Decision Protocol

When the user does not answer or requests autonomous execution:

| Decision | Phase | Default |
| :--- | :--- | :--- |
| Grains to support | 2 | All grains from `manifest.grains` where source has the required column |
| Metrics to build | 1 | All named in the intake source (issue body or definitions SQL) for the tool |
| Join strategy | 3 | Decision tree from lint join patterns |
| Domain placement | 2 | `manifest.domains` lookup (ask if new) |
| Discovery zone | 2 | First zone in `manifest.platform.zones` |
| Test window start date | 6 | MIN(TO_DATE(created_at)) from source (verified via CLI) |
| Display name override | 4 | Required if `.title()` produces the wrong label |
| GitHub issue creation | 7 | Open an issue if Mode B/C |

---

## Guardrails

15 non-negotiable rules across all phases.

| # | Guardrail | Phases |
| :--- | :--- | :--- |
| D1 | **Never skip the Phase 3 User Gate.** Design must be approved before engineering begins. Autonomous mode is the only exception and must be explicitly requested. Autonomous mode is triggered by including one of the following phrases: autonomously, no questions, run the full pipeline, or skip approval. It skips Phase 3 only -- all other guardrails remain in effect, including S5 (schema sign-off is never autonomous). | 3 |
| D2 | **Never skip quality gates.** Every phase with a VERIFY step must run it before proceeding. | All |
| D3 | **Never skip the memory store checkpoint.** Each checkpoint is BLOCKING -- write the file to disk before advancing to the next phase. Write after Phases 1, 2, 3, 4, 5, 6, and 7. | All |
| C1-C4 | **Lint code safety rules.** See [forge-lint](../forge-lint/SKILL.md) and [python-conventions.md](../forge-lint/references/python-conventions.md). Covers hardcoded zones, block comments in f-strings, zone normalization bypass, semicolons in comments. | 4, 5, All |
| S1 | **Never push without approval.** Git reads are autonomous. Git writes are human-gated. | 7 |
| S2 | **Never bypass git hooks.** No `--no-verify`. Fix the underlying issue. | 7 |
| S3 | **Never deploy to the next zone if the current zone failed.** Sequential zone dispatch with stop-on-failure per `manifest.platform.deployment_order`. | 7 |
| S4 | **Never create a PR with stale documentation.** Documentation must be current before PR creation. | 7 |
| S5 | **Never execute schema runner jobs without explicit user sign-off.** Display the full SQL content of every schema file and require explicit approval before forge-dbx is called. This guardrail applies even in autonomous mode -- schema execution is never autonomous. | 7 |
| O1 | **Never proceed without lint or dbx.** These are blocking dependencies. | 1+ |
| O2 | **Never interleave tools in batch processing.** Complete all 7 phases for one tool before starting the next. | All |
| E1 | **Never carry issue field values into Phase 3 SQL without Phase 2 evidence.** Column names, filter values, table paths, and metric name proposals from a GitHub issue or spec are hypotheses. Only values confirmed by Phase 2 live queries may appear in Phase 3 SQL expressions or be referenced in pipeline code. | 1, 3 |

---

## Edge Cases and Error Recovery

See [gate-recovery.md](references/gate-recovery.md) for the Edge Cases table, Runtime Error Catalog, Self-Healing Patterns, and Data Quality Red Flags.

At every phase boundary:

1. Log the verification failure with specific details
2. Attempt auto-fix (casing, missing comment suffix, duplicate alias)
3. If auto-fix succeeds, re-verify and continue
4. If auto-fix fails, present the issue with file path, line number, expected vs. actual

Files in this skill

  • README.md2.7 KB
  • SKILL.md29.7 KB
  • references/code-templates/dbt.md7 KB
  • references/code-templates/spark_python.md15 KB
  • references/gate-deployment.md2.8 KB
  • references/gate-discovery.md8.2 KB
  • references/gate-engineering.md5.1 KB
  • references/gate-recovery.md9.8 KB
  • references/on-demand-batch.md502 B
  • references/on-demand-removal.md1.4 KB
  • references/phase-1-intake.md9.7 KB
  • references/phase-1-issue-intake.md7 KB
  • references/phase-2-discovery.md14.9 KB
  • references/phase-3-design.md5.4 KB
  • references/phase-4-schema.md7.9 KB
  • references/phase-5-pipeline.md9.3 KB
  • references/phase-6-testing.md5.1 KB
  • references/phase-7-delivery.md6.6 KB
  • references/phase-7-deployment-playbook.md8.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…