
Claude Skills by grahama1970
github.com/grahama1970Use when the user asks how to build with OpenAI products or APIs, asks about Codex itself or choosing Codex surfaces, needs up-to-date official documentation with citations, help choosing the latest model for a use case, latest/current/default-model prompting guidance, or model upgrade and prompt-upgrade guidance; use OpenAI docs MCP tools for non-Codex docs questions, use the Codex manual helper first for broad Codex self-knowledge, and restrict fallback browsing to official OpenAI domains.
Canonical map and shared contracts for the agent-governance ecosystem: the pi.receipt_envelope.v1 boundary envelope, the component graph, and the rules for which component owns which schema. Use when wiring a skill or extension into the shared receipt world, when asking how shame, triage-error, tau, ask, project-watchdog, ops-herdr, ponytail, and Memory fit together, or when validating an envelope.
File-based inter-agent messaging with headless dispatch. Check inbox, send bugs/requests to other projects, automatically spawn headless agents to fix bugs, and track progress via task-monitor.
Artifact-driven status surfaces for long-running project-agent work. Maintains status.json, events.jsonl, proof manifests, and a stale-aware STATUS.html so humans can tell where the agent is, what passed, what is still unproven, and what decision or action is next — without dashboard theater.
Agentic evaluation of skills using multi-trial fixtures, deterministic command assertions, trajectory checks, safety constraints, and evidence-backed readiness scoring. Use when users ask for agentic evals, multi-trial skill evaluation, skill trajectory validation, or readiness scoring for a skill workflow.
Generate and query the centralized agent identity registry. Scans .pi/agents/*/AGENTS.md, parses frontmatter, outputs agents-registry.json and optionally syncs to /memory for semantic search.
Round-based context alignment before execution. Use when the human, project agent, WebGPT, scillm, ask, dogpile, memory, or project-knowledge may each hold different facts about a task; especially before ambiguous design, infographic, product workflow, high-stakes implementation, plan-iterate, project-infographic, or multi-review work.
Flexible data science analytics for any dataset. Auto-discovers schema, recommends charts, exports to create-figure. Works with JSONL, JSON, CSV from any source.
Evaluate generated Chatterbox voice files as voice-quality artifacts: affect match, arousal/valence proxies, pause placement, intelligibility inputs, clipping, loudness, and discontinuity flags. Use when reviewing Chatterbox emotional tags, pauses, Turbo/base affect delivery, Persona Dream utterance renders, or whether generated speech matches an intended product-facing affect.
Reverse-engineer features from ELF binaries. Extracts CLI commands, state machines, protocols, Zod schemas, and data models. Automatically generates a /create-walkthrough prosecution brief with Mermaid diagrams. Uses /treesitter for AST analysis of bundled JS/TS source.
Reverse-lookup glossary that turns a vague description of a web animation or motion effect into its exact term ("the bouncy thing when a popover opens" → Pop in; "the iOS rubber-band scroll" → Rubber-banding). Use when the user asks "what's it called when…", or describes a motion effect without knowing its name and wants the right word to prompt an AI or designer with. For naming an effect, not designing or building one.
Anonymize supported CSV, JSON, UTF-8 text, and SQLite files using an explicit policy through the oai-trial project. Use for anonymize data, pseudonymize exports, redact policy literals, or discover and explicitly approve fuzzy name aliases. The skill is a thin CLI/Docker interface, not another engine.
Heavy-duty "No-Vibes" debugging and hardening orchestrator. Use this for complex, stubborn bugs where `review-code` has failed, or for "Red Teaming" (hardening) a codebase. Runs multiple agents in parallel (Thunderdome) using git worktree isolation.
Apple's approach to interface design and fluid, physical motion, translated for the web. Use when building or reviewing gesture-driven UI, spring animations, drag/swipe/sheet interactions, momentum and interruptible transitions, translucent materials and depth, typography (optical sizing, tracking, leading), reduced-motion, or the design foundations (feedback, spatial consistency, restraint) behind Apple-style interfaces.
Multi-persona structured debate orchestrator. Personas research via /dogpile, consult colleagues via /ask, and argue toward nuanced synthesis on complex questions.
Search arXiv for papers and extract knowledge into memory. Use `search` to find papers, `learn` to extract knowledge.
Use when the user asks to query project memory, ask an oracle, use supported browser-backed reviewers, run Tau roundtable/single-handler workflows, ask Pi-native subagents from within Pi, run persona/deep-review workflows, generate image prompts, check OS/project health through composed skills, or run an ask DAG. This skill is the executable /ask runtime; do not replace it with an informal subagent, plain web search, or hand-written review; inside Pi, explicit Pi-native subagent targets are r...
Step back and critically reassess project state. Use when asked to "assess", "step back", "fresh eyes", "check alignment", "sanity check", "health check", "prune documentation", or "evaluate what's working". Offers documentation pruning and doc-code alignment analysis. Offer to run after major changes (don't auto-run).
Self-improvement workbench for /assistant. All the tools needed to diagnose, train, evaluate, and promote models in a continuous loop. The "warm pond" where /assistant evolves its own inference stack.
Shared GPT + classifier inference gateway for persona monitor tasks. Routes validation and classification through a 4-tier cascade: heuristic → classifier → local GPT → scillm.
Pre-flight validation and quality gates for batch LLM operations. ACTUALLY tests samples through LLM before burning tokens. Uses SPARTA contracts for DuckDB validation queries. Integrates with task-monitor for enforced quality gates.
Generate post-run analysis reports for batch processing jobs. Analyzes manifests, timings, and failures to produce comprehensive markdown reports. Optionally sends to agent-inbox for cross-project communication.
Red vs Blue team security competition orchestrator. Runs long-running overnight battles with 1000s of interactions, scoring, and insight generation.
Standardized compliance QRA benchmarks against candidate LLMs
Non-negotiable agent behavior rules. Covers: no silent failures, no error bypassing, no raw AQL, no direct imports, no parallel infrastructure, no swallowed exceptions, use existing skills, fix errors don't dodge them, transparency in verdicts, no simulated reviews, and evidence-gated decisions.
Repo-specific ArangoDB best practices: leverage text_en analyzer (stop words, stemming, BM25), use AQL functions (LEVENSHTEIN_DISTANCE, TOKENS, NGRAM_SIMILARITY, COSINE_SIMILARITY), store domain knowledge in collections not Python code, and never duplicate DB capabilities.
Advisory-first, evidence-grounded art direction and review rules for creating digital experiences that feel genuinely custom to one brand instead of template-derived. Use when a user asks for bespoke web design, a distinctive visual world, personality-led art direction, an analysis of what makes a designer's work unique, or an audit of whether a website is memorable without copying another designer's signature.
Best practices for designing, reviewing, and implementing operator chat, evidence chat, run-card chat, artifact-inspector chat, and compliance-review chat surfaces. Use when users ask for chat UX, operator console UX, agent run UX, evidence receipts, trace cards, artifact drawers, progressive disclosure, or dashboard-drift prevention.
Keep Sparta Chat usable as a modern chat interface while preserving evidence-gated compliance semantics. Use when designing, reviewing, or implementing ChatWell, InlineEvidenceCase, EvidenceWorkspace, ArtifactPanel, distance modes, voice/qid interactions, evidence receipts, artifact previews, or assistant answer ordering.
Best practices for Chatterbox or Chatterbox-Turbo voice agents, especially interruptible Embry-style agents that run memory, search, LLM, or other long-running skills concurrently. Use when designing, reviewing, or coding a voice coordinator, async task batch, JSON event stream, cancellable TTS queue, Chatterbox emotion/tag policy, pause policy, barge-in behavior, or spoken progress/preamble strategy.
Best practices for using Chatterbox and Chatterbox Turbo as an emotional voice renderer: memory-grounded utterance planning, native paralinguistic tags, spaced ellipsis pauses, render_chunks pause_after_ms, intensity/valence routing, reference-audio selection, and retained analyzer evals. Use when coding or reviewing Chatterbox answer_text, emotional tags, tone, pause policy, Persona Dream speech, or non-robotic Embry voice quality.
Cinematography grammar for AI-generated video scenes: dialogue coverage (shot/reverse-shot, 180-degree rule, eyeline direction), framing vocabulary, motivated lighting, lens/equipment language, and the AI-specific failure modes that break scenes (over-shoulder shots spawning extra bodies, eyeline flips, unmotivated light). Use when composing or reviewing shot prompts for Kling/Wan/any video model, when planning storyboard coverage for a dialogue scene, when a generated shot violates film gram...
Best practices for leading Ask compete and bakeoff workflows. Use when a user asks for competing models, isolated candidate implementations, winner selection, feature harvesting, approach comparison, model bakeoffs, or a creator competition where $ask should route browser and API handlers through Tau and the project agent must judge results against local evidence.
Best practices for conversational response behavior in voice-first agents: conversation tone, emotional steering, paralinguistic cue injection such as [laughter], wait/delay handling, interruption handling, identity-aware memory grounding, and Chatterbox-ready utterance policy. Use when designing, reviewing, or coding Embry-style chat and voice conversation behavior.
Automated COTS defense UX compliance scanner. Tests against WCAG 2.1 AA, Section 508, MIL-STD-1472H, and NIST 800-53 UI controls via CDP interaction + VLM visual analysis.
D3.js visualization best practices for performant, responsive, accessible data visualizations. Covers data joins, scales, axes, transitions, responsive SVG, interaction patterns, and accessibility. Use when writing, reviewing, or refactoring D3 visualizations.
Delivery-proof discipline for agents driving external effects: browser submits, pane messages, file writes, pushes, API calls. Use when an agent is about to claim something was sent, submitted, delivered, landed, or running; when a transport reports success but the destination shows nothing; when an agent is retrying the same failing operation; or when the human says the agent is lying, skimming, or thrashing. Encodes the 2026-08-18 session in which five browser submits failed and the agent t...
Product UX and design-practice guardrails for classifying design work, selecting applicable best-practices-* skills, preventing dashboard theater, defining mockup-first acceptance criteria, and deciding when to involve memory, dogpile, ask/scillm reviewers, interview, ux-lab, review-design, D3, React, infographic, or polish-specific skills. Use before creating, reviewing, accepting, or implementing UX direction, mockups, workflow surfaces, visual explanations, dashboards, or research-backed d...
Standards for feature-level explainability packages: strict JSONL explainers, teleprompter notes, source/debugger stops, Excalidraw/SVG diagram bindings, common architecture interview questions, and end-of-work agent transparency.
FastAPI best practices for Python control-plane services, async eval APIs, Pydantic contracts, OpenAPI surfaces, security dependencies, deployment handoffs, and Flask fallback adapters. Use when building, reviewing, or converting FastAPI/Flask services, especially skill-backed cyber-safety eval control planes.
Evidence-first typography and font-system guidance for digital products, websites, portfolios, dashboards, and design systems. Use when choosing or changing fonts, pairing typefaces, auditing overused font warnings, creating meaningful typographic hierarchy, validating font loading/provenance, or mapping typography to a visual world from best-practices-bespoke-design.
Best practices for agent-resolved GitHub tickets, including bugs, feature requests, optimizations, maintenance, questions, and triage: filing contracts, route and subagent metadata, resolver leases, deterministic verification, review evidence, WebGPT escalation, and proof-based closure.
Use for greenfield collaboration where the final product, architecture, workflow, schema, prompt, memory representation, visual direction, evaluation method, or implementation path is not fully known. Forces one named artifact, one visible contract, one candidate, one inspection, one status, and one next legal move instead of issue-tracker drift or broad architecture theater.
Repo-specific KDE/QML best practices for agentic coding: singleton design systems, property ordering, accessibility, performance (binding loops, delegate recycling), Plasma integration, and D-Bus patterns.
Create, audit, and package AI-video-ready reference packs for Kling-style element binding. Use when users ask for Kling contact sheets, Kling-ready assets, element reference packs, character reference sheets, prop sheets, scene sheets, or consistent AI-video references.
Best practices for generating faithful Kling videos via fal.ai: endpoint selection (reference-to-video vs image-to-video vs text-to-video), character identity locking with elements, prompt condensation from large pipeline artifacts (storyboards, character bibles, look locks) into Kling-sized prompts, and typed request validation. Use when Kling output does not match contact sheets or reference images, when characters render as wrong people, when choosing a Kling endpoint, when compiling a per...
Best practices for Ask one-shot runs: the same question to N seats concurrently, answers returned per seat with no consensus, no judge, and no quorum. Use when a user asks several models one question and wants to read each answer, when partial answers are still useful, or when deciding whether a request is a one-shot, a roundtable, or a competition.
Canonical rubric for evaluating, scoring, ranking, and gating career and consulting opportunities for Graham Anderson, so the "top opportunities" selection is principled and repeatable rather than bespoke per run. Use when discovering, evaluating, reviewing, ranking, or deciding which opportunities to apply to in monitor-opportunities, or when building or reviewing the opportunity-evaluator / opportunity-evaluation-reviewer subagents or the ranker. The evaluator scores AGAINST this rubric; th...
Best practices for building, reviewing, and shipping Pi TypeScript extensions and extension packages. Use when authoring Pi extensions, custom tools, lifecycle hooks, intercom bridges, status guards, provider adapters, TUI components, package manifests, or retained extension evals.
Evidence-backed standards for creating, reviewing, and hardening Pi extensions. Use when building or changing ~/.pi/agent/extensions or .pi/extensions code, package-style Pi extensions, event handlers, custom tools, TypeBox schemas, input/message/tool hooks, retry guards, final-report guards, or humorous but serious enforcement extensions such as lazy-report-shame-shame-shame.