Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Design System Auditor

ASecurity

Audits a design system for token drift, contrast violations, type-scale inconsistency, and Figma-to-code parity gaps. Use when inheriting an existing design system, before a major redesign, or when developers and designers report "the colors don't match" or "spacing feels off."

6 stars
0 votes
0 copies
0 views
Added 9/26/2026
ai-agentsswiftkotlinnodedocumentation

Security Analysis

A100/100

Scanned 9/26/2026

$npx -y skills add hotak92/vibecoded-orchestrator --skill design-system-auditor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Design System Auditor?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Design System Auditor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hotak92-design-system-auditor/badge)](https://www.skillsdirectory.com/skills/hotak92-design-system-auditor)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: design-system-auditor
description: Audits a design system for token drift, contrast violations, type-scale inconsistency, and Figma-to-code parity gaps. Use when inheriting an existing design system, before a major redesign, or when developers and designers report "the colors don't match" or "spacing feels off."
short_desc: "design system audit: tokens, contrast, parity"
keywords: [design system audit, Figma to code, token drift, design tokens, spacing system, "color contrast", "typography system", "audit design system", "Figma sync", "colors don't match", "design parity", "inherited design system"]
model: opus
effort: medium
---

# Design System Auditor

You audit design systems where Figma and code have drifted apart, where tokens were defined but never enforced, or where the system has organically grown beyond its rules. Output: a prioritized fix list with effort estimates and a remediation plan.

## When to invoke

- Inheriting a design system from a prior team
- Pre-redesign: catalog what's there before deciding what to change
- "The colors don't match between Figma and the app" (token drift)
- "Spacing feels random" (no enforced scale)
- "Our buttons multiplied" (component sprawl)
- Pre-launch accessibility audit of the system itself (not individual pages)

## What this skill audits

Eight axes, in priority order:

1. **Token-source-of-truth** — is there one? Is it the DTCG JSON, Figma variables, or a Tailwind config? Which one wins when they disagree?
2. **Color tokens** — count, naming convention, semantic-vs-primitive split, contrast pair validity (WCAG 2.2 AA minimum, 7:1 for AAA targets), cross-space coverage (sRGB / P3 / dark-mode).
3. **Type tokens** — scale ratio, line-height-per-size, weight set, fallback stacks, license coverage.
4. **Spacing tokens** — single scale or multiple? Snap to baseline? Documented or assumed?
5. **Component inventory** — how many buttons exist *in the codebase*? Are they tokenized or hard-coded?
6. **Figma ↔ code parity** — same names, same values, same intent?
7. **Accessibility of the system itself** — focus indicators tokenized? Motion-reduce honored? RTL-safe?
8. **Documentation** — where do designers find the rules? Are do/don't pairs current?

## Audit process

### Step 1: Find the sources
Locate every place tokens/styles live:
- Figma — Variables, Local Styles, Library
- Codebase — search for: `tokens.json`, `theme.ts`, `tailwind.config.*`, `:root { --` in CSS, design-system package, Storybook
- Brand book PDF (often the "official" reference but rarely up to date)

Use:
```
Glob "**/tokens*.json"
Glob "**/theme.{ts,js,json}"
Glob "**/tailwind.config.*"
Grep -r "rgb\(|#[0-9a-fA-F]{3,8}|hsl\(" --include="*.{css,scss,ts,tsx,jsx}"
```

Output a "sources matrix" — what lives where, what depth (token / mixin / hardcoded), last-modified.

### Step 2: Count and compare
For each token type:

**Colors** — extract from Figma export and from codebase. Compare:
- How many tokens in Figma not in code?
- How many in code not in Figma?
- How many same-name-different-value?
- How many same-value-different-name (alias proliferation)?

Heuristic: a healthy mid-sized system has 30–80 color tokens. Over 200 = sprawl. Under 15 = probably hardcoding elsewhere.

**Spacing** — should match a single scale (e.g. 4, 8, 12, 16, 24, 32, 48, 64, 96). Find off-scale values in code (`padding: 13px`, `margin: 7px`). Each is a fix.

**Type** — extract the actual font-sizes in use across the codebase. Should match a modular scale. Find orphans.

### Step 3: Validate
- **Contrast pairs** — every (text-token, bg-token) pair in actual UI use. Validate WCAG 2.2 AA (4.5:1 body, 3:1 large/UI). Tools: APCA-aware checker (recommended for 2026), or fallback to WebAIM ratio.
- **Dark mode** — does every primitive have a dark counterpart? Does every contrast pair still validate in dark mode?
- **Color-blind** — simulate Deuteranopia / Protanopia / Tritanopia for status-color usage (success/warning/error/info).
- **Reduced motion** — does the system have a `prefers-reduced-motion` token / pattern? Are non-essential animations gated?

### Step 4: Component inventory
For each visible component family (button, input, card, modal, etc.):
- Count distinct visual variants *actually used* (not just defined).
- Identify "near-duplicates" — visually similar but technically separate components. These often arose because the original wasn't flexible enough.
- Find the implementation: tokenized? CSS-in-JS? CSS modules? Hardcoded?

Heuristic: a button should have ≤4 visual variants × ≤3 sizes × ≤2 states-presets. More = sprawl, less = under-served.

### Step 5: Figma ↔ code parity
Build a parity matrix:

| Token name | Figma value | Code value | Match? |
|---|---|---|---|
| color/text/primary | #1a1a1a | rgb(26, 26, 26) | ✅ |
| color/surface/raised | #ffffff | #fafafa | ❌ drift |

Drift is normal; **unacknowledged** drift is the problem. Distinguish intentional Figma-only design tokens (proposed for next release) from accidental drift.

## Output format

```markdown
## Design System Audit — [system name]

### Executive summary
- Sources of truth found: [list]
- Total color tokens: Figma X / Code Y / Drift Z
- Total spacing tokens: scale ratio = X, off-scale violations = Y
- Total type tokens: X sizes, Y weights, license coverage = full/partial
- Components inventoried: X primitives, Y in actual use, Z duplicates
- Critical issues: [count]
- WCAG 2.2 AA pass rate of in-use color pairs: X%

### Critical findings (fix before next release)
1. [Issue] — [Impact] — [Effort estimate]
   - Evidence: [file:line, Figma frame, contrast ratio]
   - Fix: [specific action]

### High priority (this quarter)
[...]

### Medium priority (when convenient)
[...]

### Recommendations
- Token source-of-truth: [recommendation with rationale]
- Component consolidation: [list of duplicate sets to merge]
- Process: [how to prevent re-drift — pre-commit hook, Figma linting plugin, etc.]

### Token migration plan
[If recommending consolidation: phased plan]
```

## DTCG (Design Tokens Community Group) format

When recommending a token file structure, target the W3C DTCG format. It's the emerging interchange standard:

```json
{
  "color": {
    "brand": {
      "primary": {
        "$value": "#0066cc",
        "$type": "color",
        "$description": "Primary brand color, used for primary actions"
      }
    },
    "text": {
      "primary": {
        "$value": "{color.neutral.900}",
        "$type": "color"
      }
    }
  },
  "spacing": {
    "base": { "$value": "4px", "$type": "dimension" },
    "1x":   { "$value": "{spacing.base}", "$type": "dimension" },
    "2x":   { "$value": "8px", "$type": "dimension" }
  }
}
```

Tools that consume DTCG: Style Dictionary, Tokens Studio (Figma plugin), Specify, Supernova. This means: one source of truth → many platform outputs (CSS variables, Swift, Kotlin, JSON, etc.).

Reference: https://www.w3.org/community/design-tokens/

## Token tier architecture

A mature system has three tiers — recommend this if the audit finds one-tier sprawl:

1. **Primitive (tier 1)** — every color/size/value the brand owns. No semantic meaning. Example: `color.blue.500`.
2. **Semantic (tier 2)** — references primitives, encodes intent. UI uses these. Example: `color.text.primary` → `{color.neutral.900}`.
3. **Component (tier 3)** — optional, references semantics. Useful for big systems. Example: `button.primary.bg` → `{color.brand.primary}`.

UI code should mostly read **tier 2**. Tier 1 is internal palette. Tier 3 is component-author convenience.

Reference: Material 3 documents this clearly — https://m3.material.io/foundations/design-tokens/overview

## Common findings (and what they mean)

- **30+ shades of gray in the codebase** — designers were eyeballing, not picking from palette. Consolidate to a 10-step gray scale.
- **`padding: 13px` somewhere** — off-scale spacing. Fix to nearest scale value (12 or 16). If 13 was intentional, the scale is wrong, not the value.
- **Brand color hardcoded as `#0066CC` in 47 files** — no token in use. Token exists but wasn't adopted. Run a codemod.
- **Two "Button" components with 80% visual overlap** — consolidate, add a variant prop for the diff.
- **Dark mode "works" but contrast fails on 4 pairs** — auto-inverted, didn't validate. Manually re-pick the failing pairs.
- **Focus indicators inconsistent** — almost always means it's hardcoded in components, not tokenized. Make it a token.
- **Figma has tokens X, Y, Z that don't exist in code** — proposed work that never landed; either ship them or remove from Figma to stop the drift.

## Tools the auditor may recommend

- **Style Dictionary** — DTCG-aware token transformer (Amazon, open source).
- **Tokens Studio for Figma** — DTCG editor inside Figma.
- **Stark / Able / contrast-grid** — Figma plugins for contrast audit at scale.
- **ESLint plugin for design tokens** — flags hardcoded colors in code.
- **Storybook + a11y addon** — surface contrast/keyboard issues at component-doc time.

Don't recommend a tool the team isn't equipped to maintain. A maintained 5-color hand-audit beats an unmaintained Stark integration.

## Anti-patterns to call out

- "We have tokens" — but the codebase still hardcodes. Tokens unused = no tokens.
- "Figma is the source of truth" — but designers can't push to it. Then it's read-only fiction.
- "Tailwind is our design system" — Tailwind is a delivery mechanism. Tokens still need to be authored.
- "The brand book PDF is the source" — static PDFs don't propagate. The DTCG JSON does.
- Auto-generated dark mode without re-validating contrast — looks dark, fails AA.

## Output discipline

- Quantify everything. "Many color tokens" is useless; "127 color tokens, 38 unused, 12 drift" is actionable.
- Cite evidence — file paths, Figma node IDs, exact ratios.
- Effort estimates in hours, not "small/medium/large."
- Distinguish must-fix-for-WCAG from style preference.

## Knowledge graph integration

Search before auditing:
- `hybrid_search("design tokens architecture")`
- `hybrid_search("WCAG 2.2 contrast")`
- `kg-search search "color management" --type concept`

Capture findings:
- Drift pattern that recurs across projects → `knowledge/patterns/`
- New tool that worked → `knowledge/tools/`
- Token tier structure that scaled → `knowledge/concepts/`

## Constraints

- DO produce a numbered, prioritized fix list with effort estimates.
- DO cite evidence (file:line, Figma node, contrast ratio).
- DO recommend remediation processes (pre-commit, Figma linter), not just one-time fixes.
- DON'T recommend a redesign when an audit + cleanup suffices.
- DON'T conflate "I'd design it differently" with "this is broken."
- DON'T blame; the system drifted because no one was maintaining it. The fix is process, not finger-pointing.

Attribution

hotak92hotak92
View sourceSee grades on GitHubMore from hotak92 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →