
Claude Skills by mcorbett51090
github.com/mcorbett51090Design an accessible-by-default component pattern, semantic HTML first and ARIA only where needed. Reach for this on a design-system or component question.
Rank audit issues by user-impact and effort into a sequenced remediation plan with owners. Reach for this when there are more fixes than time.
Audit a page set against a named WCAG version and level, classify issues by severity/level, and compute a weighted conformance score. Reach for this on a conformance question.
Verify keyboard operability and screen-reader parity hands-on with the assistive technology real users use. Reach for this on a parity question.
Compute the WCAG contrast ratio from hex foreground/background values and check AA/AAA for normal and large text. Reach for this on any color question.
Audit segregation of duties and chart-of-accounts hygiene before trusting the books. Reach for this on a controls or data-quality question.
Estimate bad-debt from AR aging buckets weighted by loss rate. Reach for this on a receivables-risk question.
Read the cash conversion cycle (DSO + DIO − DPO) and locate trapped or surrendered cash. Reach for this on a cash question.
Reconcile bank and balance-sheet accounts to source before any statement ships. Reach for this first on any reporting question.
Run the period-end close on a cadence: critical-path checklist, days-to-close, bottleneck. Reach for this on a close question.
Design the tools/functions an agent calls and the context/memory strategy that keeps it coherent — unambiguous tool names, typed parameters, examples-in-description, errors that teach recovery, a small-enough tool count, plus a context plan (what stays in-window, what gets summarized, what moves to external memory) and short-term-vs-long-term memory design under a per-turn token budget. Reach for this when the user says 'the agent keeps calling tools wrong', 'the context overflows', or 'the a...
Evaluate an agent and harden its failure modes before it touches real traffic — an offline eval harness scored on task-completion, trajectory, and tool-use correctness against a fixed task set, plus loop hardening (step/tool-call caps, timeouts + retries, stop conditions, human-in-the-loop on irreversible actions) and tracing so every step/tool-call/token-cost is observable, with cost and latency reported alongside quality. Reach for this when the user asks 'how do I know this agent works?', ...
Decide whether a task should be an agent at all — and if so, single-agent vs multi-agent, the orchestration topology, and the framework — by traversing the agentic decision tree (agent-vs-workflow gate → single-vs-multi → topology → framework), returning a go/no-go verdict that defaults to 'a fixed workflow or a single LLM call wins' unless the control flow is genuinely unknowable in advance. Reach for this when the user asks 'should we build an agent for this?', 'do we need multiple agents?'...
Calibrate the OpenAI Codex reasoning-level dial before recommending a model upgrade. Maps task type, failure mode, and budget to the right reasoning effort level — ensuring developers exhaust the reasoning dial on the current model before paying for a bigger SKU. Domain-specific to the Codex reasoning API.
Scope an autonomous AI coding agent task so it is well-bounded, recoverable, and matched to the right model tier before it runs. Reach for this skill before any long, unsupervised, or multi-step agentic run — a poorly scoped task on a frontier model is both expensive and hard to debug.
Estimate a coding task's context demand and match it to a model tier whose window is sufficient, without quoting specific token counts that churn monthly. Reach for this skill when the task involves large codebases, long conversation histories, or multi-file agentic runs where context overflow would produce silent truncation errors.
Audit a developer's GitHub Copilot configuration across all surfaces (completions, chat IDE, coding agent, cloud agent, mobile) to identify model gaps, plan mismatches, and org-policy constraints. Reach for this skill before recommending a Copilot model — surface and plan scope the answer, not just the model name.
Check whether a Grok model ID in active use has been retired or silently redirected, with special attention to billing consequences. Reach for this skill before any Grok recommendation and whenever a developer mentions a specific Grok model ID — retirement redirects can incur unexpected charges at the new model's pricing.
Check the ai-coding-model-guidance knowledge bank for stale entries and produce a prioritized refresh list. Reach for this skill when the knowledge bank's retrieval date is more than 4 weeks old, when a consumer reports a model discrepancy, or on the monthly researcher-reminder cadence.
Compare model options across GitHub Copilot, OpenAI Codex, and xAI Grok for a single task when the developer has not committed to one ecosystem. Produces a structured side-by-side that surfaces availability, tier mapping, and key tradeoffs without naming specific volatile numbers.
Audit an organization's AI coding tool model-access policies across GitHub Copilot Business/Enterprise, OpenAI org-level controls, and xAI API governance. Reach for this skill when an enterprise team reports unexpected model access, when a compliance review requires documenting which models the org has allowed or blocked, or before rolling out a new model to a large org.
Hard gate when a coding agent/surface returns quota, rate-limit, weekly/monthly cap, spend-limit, or tokens-exhausted. Never silent-fail or blind-retry — try param/effort/scope levers before vendor failover, then traverse the vendor-neutral tier tree before naming a substitute SKU. Deep layer: coding-agent-levers-playbook.
Compute cost per request and right-size the context to fewest-high-precision chunks. Reach for this on a cost/context question.
Build a judgment set and measure recall@k, precision@k, faithfulness, and answer-relevance with a baseline. Reach for this before shipping any change.
Separate retrieval failure from generation failure by measuring recall@k before touching the model. Reach for this first on wrong answers.
Add citations, refuse-on-empty-retrieval, and context-constraint to cut hallucination. Reach for this on a faithfulness question.
Tune chunk size, overlap, and structure-awareness against the eval and the context budget. Reach for this on a chunking question.
Scope an AI red-team engagement by threat-modeling the system (assets, attackers, trust boundaries), splitting safety from security, traversing the attack-taxonomy decision tree to a prioritized OWASP LLM Top 10 / MITRE ATLAS attack list, and setting the rules of engagement plus likelihood×impact success and severity criteria. Reach for this when the user asks "how should we red-team this LLM feature?", "what should we attack first?", "is this a safety or a security problem?", or "what are th...
Triage red-team findings by likelihood×impact and drive defense-in-depth remediation — layered input/output guardrails, injection-resistant prompt structure, least-privilege tool scoping, allow-lists, human-in-the-loop on high-impact actions, output-handling hygiene, and rate/cost limits — then retest each fix with the exact attack that found it and bake it into the regression harness. Reach for this when the user asks "we have a pile of red-team findings — what do we fix and how?", "harden o...
Execute the prioritized attacks against an AI system within the rules of engagement — direct and indirect prompt injection, jailbreaks (roleplay, encoding, many-shot, crescendo), training-data extraction and data exfiltration, agentic tool-abuse / excessive agency, and multimodal attacks — capturing each as a reproducible payload plus transcript, then automating what repeats into a PyRIT / Garak / Promptfoo red-team / Giskard harness with a scorer and a CI regression gate. Reach for this when...
Keep the warehouse trustworthy: dbt tests (not_null/unique/accepted_values/relationships) gating the build in CI, source-freshness checks, model contracts at consumer boundaries, singular tests for business invariants, and anomaly detection beyond schema tests.
Design a dbt CI pipeline that gates every pull request: compile, run, test, and check source freshness in an isolated developer schema; enforce model contracts on published marts; run slim CI on changed models only using dbt state comparison; and block merges on test failures or contract violations.
Model in dbt across staging -> intermediate -> marts layers, choose materialization (view/table/incremental) by the trade, write correct incremental models (reliable unique key, is_incremental filter, late-data strategy), and keep it DRY with refs/sources/macros.
Build reliable incremental dbt models: choose the right unique_key and strategy (append, merge, delete+insert), handle late-arriving data and out-of-order events, write a safe is_incremental filter, and design the full-refresh fallback — so the model is idempotent from day one.
Build a governed semantic/metrics layer: define each metric once as metrics-as-code (dbt Semantic Layer/MetricFlow) with explicit grain and filters, model entities/dimensions to prevent fan-out, and expose one contract every BI tool consumes — ending metric drift.
Step-by-step playbook for deprecating and sunsetting an API version — header strategy, consumer communication timeline, traffic monitoring gates, and the SDKs/portal update checklist.
Playbook for designing cursor (keyset) pagination on list endpoints: cursor encoding, response envelope, query parameter contract, and the migration path away from offset pagination.
Playbook for designing safe-to-retry POST and PATCH operations using Idempotency-Key — covers key format, dedup window, stored-response replay, conflict handling, and OpenAPI declaration.
Step-by-step playbook for writing a contract-first OpenAPI 3.1 document — from info block to path items, component reuse, and Spectral pre-flight. Covers resource modeling, status codes, error schema, and pagination shape.
Playbook for designing a consistent RFC 9457 Problem Details error model — type URIs, extension members, status code mapping, and a catalog template. Prevents per-endpoint bespoke error shapes.
Playbook for writing and wiring a Spectral ruleset that enforces an API style guide in CI — covers rule anatomy, severity levels, custom functions, and the recommended core-rules baseline.
Pick the right hypothesis test for a described scenario by traversing the test-selection decision tree (data type → #groups → paired? → assumption gate → test), then return the recommended test, its assumption checks, its nonparametric fallback, and a ≤10-line runnable snippet. Reach for this when the user asks "which test do I use?" or hands over two-or-more groups/variables to compare. Used by `applied-statistician` (primary).
Analyze a completed A/B test or experiment defensibly — check it against the pre-registered plan, run the primary-metric test, report effect size + CI (not just p), check guardrail metrics, apply a multiple-comparison correction across metrics/segments, and screen for the peeking/p-hacking pitfalls before declaring a winner. Used by `applied-statistician` (primary).
Compute or advise the sample size an experiment needs BEFORE it launches — from α (0.05), power (0.80), and a minimum detectable effect (MDE) — or compute the power/MDE a fixed sample can achieve. Prevents the underpowered-study pitfall and is the prerequisite to any A/B test. Returns the n, the assumptions behind it, and a runnable snippet. Used by `applied-statistician` (primary).
Review or design a regression model or a time-series forecast so it's defensible — pick the model family (OLS / logistic / Poisson GLM; ARIMA / SARIMAX / ETS), check the assumptions that matter for that family, report honest prediction/confidence intervals, and screen for overfitting, data leakage, and "coefficient = cause" overreach. Used by `applied-statistician` (primary).
Decide whether a dashboard metric movement, comparison, or trend is signal or noise — and annotate it honestly (significance, confidence interval, "not enough data yet"). The interop seam with data-platform — invoked by `data-platform/dashboard-builder` when a widget shows a comparison/trend that needs a statistical-validity annotation. data-platform answers "is this number correct?"; this skill answers "is it real?". Used by `applied-statistician` (primary) + `data-platform/dashboard-builder`.
Treat XR comfort, physical safety, and accessibility as requirements: hold a sustained framerate, choose comfortable locomotion, design for the tracking volume / guardian / play-space and passthrough so users don't hit walls, and ship accessibility options (seated mode, one-handed paths, snap turn, captions, adjustable text) as defaults. Comfort research verify-at-use.
Hold the XR frame budget: derive the per-eye ms target from the device refresh rate, profile on device to find the CPU-vs-GPU-vs-thermal bound, cut draw calls / overdraw / fill rate in order, and use foveated rendering and reprojection as headroom not a crutch — budgeting for the thermal-sustained clock, not peak. Device numbers verify-at-use.
Design and implement XR interaction: hand / controller / gaze input on an OpenXR action abstraction, locomotion chosen for comfort (teleport / dash / snap-turn vs smooth + vignette), reachable and readable 3D UI, and intentional grab/physics — with accessibility built in as a requirement. Input-API specifics verify-at-use.
Choose the XR target platform (standalone headset vs PC-VR vs WebXR vs mobile-AR) and engine (Unity / Unreal / native-OpenXR / WebXR) on the use-case, audience device, and distribution channel — then commit to an OpenXR-first architecture and derive the per-eye perf budget the target implies. Device/version specifics verify-at-use.