
Claude Skills by mcorbett51090
github.com/mcorbett51090Connector-specific configuration patterns for ELT pipelines — QuickBooks OAuth + rate-limit handling, Stripe webhook + batch hybrid, Salesforce Bulk API 2.0, HubSpot API v3, GA4 BigQuery export, Shopify GraphQL Admin API. Used by `etl-pipeline-engineer` when configuring an Airbyte / Fivetran / n8n connector against a real source.
Use when stitching the same real-world entity — a customer account — across multiple source systems into one conformed spine; e.g. Salesforce + Planhat + Intercom + Slack. Operationalizes the deterministic-keys-before-fuzzy best practice: inventory candidate keys, build the precedence ladder, construct bridge_account_xref with confidence + match_method, quarantine unresolved records, run resolution_audit, and gate on stewardship review for low-confidence matches. K-12 LEAID confidence-tier la...
Scaffold Cube semantic-layer schemas with mandatory `securityContext` baked in for multi-tenant customer-facing dashboards. Includes measure/dimension patterns, pre-aggregation hints, view-level partner-facing query surface, and the cross-boundary denial test. Used by `dashboard-builder` on Case-C productized-SaaS engagements.
Systematically walk a dashboard page by page and judge whether its structure makes sense, whether it tells a coherent story, and whether it guides the user toward action — for a new build's final gate and for hardening/upgrading an existing dashboard. Invoked by `dashboard-builder` (primary) and directly for a standalone hardening request.
Tune interactive-dashboard performance against a per-widget-class budget — Cube pre-aggregation design, Postgres / DuckDB materialized views, cache layers (Cube + Redis + browser TanStack Query), the per-widget profile loop (measure → identify the slow stage → fix at the lowest-cost layer). Reach for this skill when a dashboard exceeds the 1-2s widget target, or proactively before adding a heavy widget. Used by `dashboard-builder` (primary).
Design data-quality tests that catch real bugs — uniqueness / not-null / referential integrity / freshness / row-count drift / value-range / cross-source reconciliation. dbt-test mechanics for each, severity tiers (error vs warn), the runbook-entry-per-failing-test discipline, and when to escalate to Great Expectations or Monte Carlo / Bigeye. Reach for this skill when launching a new pipeline OR after a "the numbers were wrong" incident. Used by `etl-pipeline-engineer` (primary).
Scaffold a dbt project that ships — sources → staging → intermediate → marts → metrics layer discipline, generic + custom tests, doc-blocks for every model, exposure tracking, RLS-safe role separation (build-role vs query-role), CI shape (`dbt build` on PR), and a dev/prod env-promotion shape. Reach for this skill at engagement start (greenfield warehouse) or when a dbt project has decayed into ungoverned models. Used by `etl-pipeline-engineer` (primary) + `dashboard-builder`.
Configure CSP `frame-ancestors`, iframe `sandbox` attributes, postMessage origin checks, and web-component shadow-DOM boundaries for embedded dashboards. Invoked by `ravenclaude-core/security-reviewer` when a diff touches embed-auth flow; generated alongside dashboard code by `dashboard-builder`.
Canonical 2026 JWT-embed flow for dashboard embedding — required claims (`sub`, `tenant_id`, `iat`, `exp`, `iss`, `aud`, `nonce`), tool-specific verification (Superset guest tokens, Metabase JWT URLs, Cube Authorization Bearer, Power BI MSAL-via-AAD), 5-15 min expiration policy, cross-boundary denial test contract. Invoked by `ravenclaude-core/security-reviewer` for any embed-auth review.
Migrate a single-tenant database / dashboard stack to multi-tenant — `tenant_id` column propagation + backfill, RLS / semantic-layer scope rules introduced post-hoc, JWT-claim shape migration, embed-token cutover plan with parallel + backout window, and the mandatory cross-boundary denial test before flipping the switch. Reach for this skill when an engagement shifts from one-client deliverable to productized SaaS. Used by `database-setup-guide` (primary) + `dashboard-builder`.
Author Postgres Row Level Security policies (force-on, CI-deployed, denial-tested) for multi-tenant dashboards. Encodes the closeness-to-data invariant — tenant isolation at the closest-to-data layer the viewer's token cannot influence. Covers Postgres RLS + semantic-layer enforcement (Cube `securityContext`, Power BI DAX roles, Fabric, Snowflake row-access). Invoked by `ravenclaude-core/security-reviewer`.
Select a dashboard-engagement stack via the Case A/B/C/D Mermaid decision tree (Portfolio / Per-client / Productized SaaS / Pipes-only) — surfaces the per-viewer-pricing-trap heuristic, recognizes the EdTech LMS connector-gap, returns a populated `stack-decision-record.md`. Invoked by `ravenclaude-core/architect` via inline prior.
Take vendor-specific raw support-ticket tables in the warehouse (Zendesk, Freshdesk, Intercom Conversations + Tickets, SFDC Case, JSM, HubSpot, Help Scout, Front) and produce conformed `fct_ticket` + `fct_conversation_event` models with SFDC Account / Planhat Company bridge resolution. Includes the per-vendor field-mapping table, status / priority / channel conformance maps, escalation-flag derivation rules, SLA-breach derivation, the connector-method decision tree, and the dbt-test shape tha...
Pick the right data-quality approach and tooling for a described stack by traversing the data-quality tooling decision tree (already-on-dbt? → check shape, known-rule vs unknown-over-time → build vs buy → where checks run → tool), then return the recommended contracts/tests/monitors mix, the tool (dbt tests / dbt-expectations / Great Expectations / Soda / Elementary / a managed platform / warehouse-native), where each check runs, the block-vs-warn policy, the SLAs, and the conditions that wou...
From a dataset and its consumers, derive the producer-boundary data contract (schema, semantics, freshness and volume expectations, ownership) and the concrete validation test suite (not-null, unique, accepted-values, referential integrity, distribution/value-range), each test carrying a severity. Reach for this when the user asks "write the data contract for this table", "what tests should this dataset have?", or "define the guarantees this producer owes its consumers". Used by `data-quality...
Stand up the observability monitors that watch for the unknown over time — freshness, volume, schema-drift, and distribution/anomaly monitors — each anchored to a baseline and a tolerance (not a hard-coded magic number), then wire owner-routed alerting and link a data-incident runbook. Reach for this when the user asks "set up freshness/volume monitors", "alert us before a stakeholder notices bad data", or "detect schema drift and distribution anomalies". Used by `data-quality-engineer` (prim...
Run a disciplined exploratory pass before anyone models: profile the data (shape, types, missingness, cardinality, distributions), make and document cleaning decisions, visualize distributions and relationships, spot leakage candidates and target-definition problems, generate hypotheses, and communicate findings with their uncertainty.
Engineer features and fit, select, and honestly evaluate classical models on tabular data: build leakage-aware features, set a baseline, choose between regression / trees / gradient boosting, build a leakage-safe cross-validation harness with every transform fit inside the fold, and pick the metric from the decision and its cost structure.
Make an analysis re-runnable by anyone: enforce notebook hygiene (top-to-bottom-clean, no out-of-order state), pin environments and dependencies in a lockfile, set and thread random seeds, version data and artifacts, track experiments (params + metrics + code/data version per run), and convert a run-once notebook into a scripted, deterministic pipeline.
Step-by-step playbook for standing up a Change Data Capture pipeline with Debezium — source connector configuration, slot management, snapshot strategy, and handling common failure modes.
Structured triage procedure for diagnosing and resolving Kafka consumer group lag — distinguishes slow consumer, under-partitioned topic, rebalance storm, and broker-side causes.
Run the streaming platform reliably: partition for the ordering you need (order is per-partition; the key is the guarantee), govern schemas with a registry + compatibility rules, configure producer durability (acks/idempotence) and consumer offsets/idempotency, and ingest via CDC not dual-writes.
Step-by-step playbook for evolving Avro/Protobuf/JSON Schema event schemas without breaking consumers, covering compatibility modes, migration patterns, and registry operations.
Playbook for managing Avro/Protobuf/JSON Schema schemas in a registry — choosing a compatibility mode, executing safe schema changes, and handling breaking changes without consumer downtime.
Process streams correctly: aggregate on event-time with watermarks (not processing-time), window deliberately (tumbling/sliding/session), handle late data explicitly, checkpoint and TTL-bound state, join with aligned time, and design for backpressure.
Decide streaming vs batch honestly by the real latency need (sub-minute reaction -> streaming; hourly/daily -> batch via data-platform), then design the topology, platform choice, and delivery semantics if streaming is justified.
Playbook for sizing and configuring a database connection pool (PgBouncer or application-level) — calculating the right pool size, choosing the pooling mode, diagnosing pool exhaustion, and avoiding the common over-provisioning trap.
Operate the database reliably: pool connections with sane limits, choose the isolation level deliberately (knowing its anomalies), keep transactions short, route replica reads carefully, and test restores — a backup is only real if you've restored it.
Tune queries with evidence: read EXPLAIN (ANALYZE, BUFFERS) first, choose the right index type and column order (B-tree/partial/composite/covering/GIN), rewrite to be sargable, kill N+1 at the SQL level, and keep statistics fresh.
Design a correct relational schema: normalize to 3NF, denormalize only with measured evidence and a named consistency cost, push constraints (PK/FK/UNIQUE/CHECK/NOT NULL) into the database, and choose precise data types.
Evolve schemas without downtime: expand/contract (add -> backfill in batches -> switch -> drop) across separate deploys, lock-aware DDL (create indexes CONCURRENTLY, nullable-add-then-validate), reversibility, and ordered versioned migrations.
Design a backup + point-in-time-recovery strategy from RPO/RTO and — critically — the restore-verification that proves it works: cadence, retention, failure-domain isolation, and a scheduled test restore measuring actual restore time. Reach for this when backups are unproven, when setting RPO/RTO, after a near-miss, or before relying on recovery. Pairs with ha-topology-and-failover.
Triage and mitigate a live production database incident — classify latency vs errors vs saturation, rank the suspects (lock contention, replication lag, connection storm, runaway query, disk-full, failover), run the per-mode diagnostic, and apply the least-blast-radius reversible mitigation first. Reach for this when a DB is slow/erroring in prod now. Pairs with zero-downtime-migration when a change caused it.
Design a production database's high-availability topology from its RPO/RTO targets — replica type (sync/async/quorum), failure-domain spread (multi-AZ/region), and the failover mechanism with a split-brain guard. Reach for this when a DB needs to survive an AZ/region outage, when failover has never been designed, or when 'we have a replica' is mistaken for HA. Pairs with backup-and-restore-verification.
Plan a schema change on a live production database with no downtime using expand-contract — additive expand, dual-write, batched/throttled backfill, cutover, then contract — with lock duration and replication lag managed. Reach for this before any non-additive change to a hot table (NOT NULL, rename, type change, drop) or a large backfill. Pairs with db-incident-triage if a migration goes wrong.
Design a Databricks lakehouse from the query and freshness SLO backward: the medallion layering (what earns bronze/silver/gold), the Delta table & partitioning strategy per layer (low-cardinality partition vs liquid clustering, OPTIMIZE/Z-ORDER/VACUUM cadence, MERGE/CDC vs append/overwrite write pattern), the batch-vs-streaming call (scheduled batch is the default; Auto Loader / Structured Streaming / DLT only for a real sub-hour SLO), the Unity Catalog governance layout (catalog/schema, mana...
Diagnose a slow or failing Databricks/Spark job from the EVIDENCE (the Spark UI stages/task-skew/shuffle-spill/GC and the query plan) rather than guessing, then apply the fix the evidence points to — AQE skew-join or key salting for a hot key, broadcast for a small-dim join, repartition/coalesce for partition sizing, OPTIMIZE/compaction for the small-file problem, and writing-instead-of-collecting for driver OOM — and bring the DBU cost down (auto-termination, jobs-vs-all-purpose compute, rig...
Read overhead as a % of collections against the ~62% median before cutting any single cost, so the diagnosis targets the real driver. Reach for this on a margin problem.
Raise treatment-plan acceptance through presentation and sequencing rather than discounting. Reach for this when big plans don't close.
Read the effective fee by plan and manage PPO write-offs as a deliberate strategy, not an accident. Reach for this when adjustments erode margin.
Read banked-vs-produced dollars and recover the collection ratio toward 98%+, so production becomes income. Reach for this when collections slip.
Read doctor and hygiene production per hour, not per day, to expose the real capacity story. Reach for this on a production question.
Design and build an accessible, composable library component with a contract-grade public API — deciding composition vs configuration and controlled vs uncontrolled, baking in roles/focus/keyboard from v1, and shipping a story and usage docs. Traverses the component-API branch of the design-systems decision tree. Reach for this when the user asks 'compose or props for this component?', 'build an accessible Menu/Dialog/Combobox', 'controlled or uncontrolled?', or 'what should this component's ...
Version a component library like the public contract it is (semver + changesets), ship breaking changes with codemods and loud deprecation, and drive adoption with metrics. Traverses the versioning branch of the design-systems decision tree: semver mapping → changesets/release flow → breaking-change/codemod policy → deprecation → adoption metrics. Reach for this when the user asks 'how do we version the library?', 'set up changesets', 'how do we ship this breaking change?', or 'how do we driv...
Structure design tokens across the primitive → semantic → component tiers, with theming/dark-mode and multi-brand solved as a semantic-tier value swap, then emit the platform outputs consumers use. Traverses the token branch of the design-systems decision tree: source of truth → tier structure → naming → theming → platform outputs. Reach for this when the user asks 'how should our tokens be structured?', 'set up design tokens', 'how do we do dark mode / multi-brand?', or 'primitive vs semanti...
Decide Electron vs Tauri vs native vs PWA by the app's real needs (bundle size, native-API depth, team skills, security surface, update needs), then set the process/security model and the renderer/backend boundary.
Harden an Electron app to the secure baseline — contextIsolation on, nodeIntegration off, sandbox on, a strict CSP, no remote module — and bridge to the OS through a narrow typed contextBridge backed by validated ipcMain handlers.
Wire native OS integration the way each platform expects — tray/menu-bar, application menus, notifications, file associations, custom-scheme deep links routed through a single-instance lock, and secrets in the OS credential store.
Package, code-sign, and notarize a desktop app for Windows (Authenticode/EV) and macOS (Developer ID + notarytool + staple), and ship a safe signed auto-update (signature verified before apply, channels, staged rollout, rollback, version floor).
Expose native capability in Tauri through #[tauri::command] handlers with validated input, authorized by a least-privilege capabilities/permissions allow-list (v2) scoped to the windows that need it — no wildcard fs/shell scopes.