Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Eval Golden Set Builder

ASecurity

Playbook for constructing a high-signal golden evaluation set: sampling strategy, input diversity, reference-answer authoring, grader pairing, and the minimum-viable size thresholds that make a delta meaningful. Owned by eval-engineer.

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentspythonrustgo

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill eval-golden-set-builder --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Eval Golden Set Builder?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Eval Golden Set Builder
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-eval-golden-set-builder/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-eval-golden-set-builder)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: eval-golden-set-builder
description: "Playbook for constructing a high-signal golden evaluation set: sampling strategy, input diversity, reference-answer authoring, grader pairing, and the minimum-viable size thresholds that make a delta meaningful. Owned by eval-engineer."
---

# Eval Golden Set Builder

## When to invoke

- Starting evals from scratch on a new prompt or agent.
- An existing golden set is producing noisy deltas (every change looks "better").
- A product change (new tool, new model, new system prompt) needs a before/after regression check.
- LLM-judge results are inconsistent across runs.

## Step 1 — Define what you are measuring

Before collecting examples, answer these:

| Question | Why it matters |
|---|---|
| What is the unit of success? | A correct final answer? A correctly-chosen tool? A well-structured JSON output? |
| Who is the judge? (human / programmatic / LLM) | Determines reference-answer format |
| What is the blast radius of a regression? | Determines minimum set size and grader strictness |

One eval, one signal. Do not try to evaluate "everything" in one set. Separate factual accuracy from tone from tool-call correctness — they need different graders.

## Step 2 — Sampling strategy

| Stratum | Target fraction | What goes here |
|---|---|---|
| Happy path | 40 % | Canonical inputs the system handles today |
| Edge / boundary | 30 % | Near-miss inputs, ambiguous phrasing, missing fields |
| Adversarial | 20 % | Injection probes, jailbreak-adjacent inputs, known failure modes |
| Regression seeds | 10 % | Any input that caused a past production bug |

Minimum viable set: **50 examples** for a binary pass/fail grader; **100+** for a scored (0–5) rubric or LLM judge (smaller sets produce deltas too noisy to trust).

## Step 3 — Authoring reference answers

1. **Human-authored ground truth first.** Write the ideal answer before you see what the model produces — post-hoc rationalization corrupts the set.
2. **For tool-call evals:** record the expected `tool_use` block — `name` + exact `input` keys. Use `"contains"` matching (not string equality) for free-text arguments.
3. **For structured-output evals:** store the expected JSON; use JSON Schema validation as the grader, not string diff.
4. **For open-ended answers:** write a rubric (3–5 criteria, each scored 1–3) that a judge (human or LLM) can apply without seeing the question first. Store it alongside the example.

```jsonl
{"id": "q_001", "input": {"messages": [...]}, "expected": {"tool": "search_docs", "query_contains": "refund policy"}, "grader": "tool_call_match"}
{"id": "q_002", "input": {"messages": [...]}, "expected": {"rubric": "criteria_accuracy_helpfulness"}, "grader": "llm_rubric"}
```

## Step 4 — Grader pairing

| Output type | Grader | Notes |
|---|---|---|
| Exact JSON / schema | JSON Schema validator | Fast, deterministic, free |
| Tool call correctness | Name + key-subset match | Allow partial `input` matches |
| Binary pass/fail factual | String `in` / regex | Hand-write assertions |
| Scored rubric / tone / style | LLM-as-judge (Haiku via Batch) | Run via Anthropic Batch at 50 % cost; randomize answer order per call to reduce position bias |

**LLM-judge discipline:** present judge with rubric + response only (not the question) to prevent question-loading bias. Run 3× and take majority vote on borderline cases.

## Step 5 — Baseline and delta protocol

```python
# Run on every prompt/model/tool change
baseline = run_eval(golden_set, old_config)
candidate = run_eval(golden_set, new_config)
delta = candidate.mean_score - baseline.mean_score
p_value = ttest(candidate.scores, baseline.scores).pvalue
print(f"delta={delta:+.3f}  p={p_value:.3f}  n={len(golden_set)}")
# Merge only if delta >= +0.02 and p < 0.05
```

Treat a delta < 0.02 on a 0–1 scale as noise, not improvement. Track per-stratum scores — an adversarial regression hidden by a happy-path improvement is still a regression.

## Pitfalls

- Seeding the golden set from model output ("label what the model said as correct") — creates a circular reference that only confirms the current model's biases.
- A set with 90 % happy-path inputs — edge and adversarial cases are where regressions actually hide.
- Comparing runs on different dates without pinning the model version and temperature (`temperature: 0` for evals).
- Re-running a failing test until it passes and calling that "fixed" — add the fixed case as a regression seed; don't delete the test.

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →