Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Synthetic Data

ASecurity

Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgoperformance

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill synthetic-data --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Synthetic Data?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Synthetic Data
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-synthetic-data/badge)](https://www.skillsdirectory.com/skills/aicodedecode-synthetic-data)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: synthetic-data
description: Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.
category: ai-research
---

## Overview

Synthetic data is training data generated by models rather than collected from
humans or the world: instruction-response pairs, reasoning traces, code with
tests, preference judgments, tool-use trajectories. It's the engine behind most
modern post-training — distillation, self-improvement loops, and domain adaptation
all run on synthetic data. Done well, it converts a strong teacher model's
capabilities into targeted training signal for a student. Done poorly, it produces
fluent garbage that teaches the student the teacher's failure modes.

The core discipline is treating synthetic generation as a data pipeline with
quality gates, not as "prompt a big model and hope." Every synthetic dataset needs:
a generation step with controlled diversity, a verification step that filters for
correctness (execution, entailment checks, cross-model agreement), and a validation
step measuring downstream effect on the student. The failure mode to fear is model
collapse — training on unfiltered synthetic data degrades quality over
generations — which is why verification and mixing with real data matter.

Synthetic data is leverage: it multiplies a strong teacher into arbitrary
quantities of targeted training signal. Like all leverage, it amplifies errors
too. The pipeline's job is to keep the signal while discarding the errors.

## When to use

- Post-training a model for a domain where human data is scarce or expensive
  (specialized reasoning, low-resource languages, niche tools).
- Distilling capabilities from a stronger teacher into a smaller, cheaper student.
- Generating instruction-tuning data at a scale humans can't match.
- Creating adversarial or edge-case examples to patch specific weaknesses found in
  evals.
- Bootstrapping preference data (pairs judged by a strong model) for DPO or
  reward modeling.
- Augmenting thin slices of real data (rare classes, edge cases) where collection
  is impractical.

## Core concepts

- **Generation with diversity control.** Naive sampling collapses to the teacher's
  modal outputs. Use temperature variation, persona/topic seeding, and explicit
  diversity prompts ("generate a problem unlike these examples") to cover the
  space.
- **Verification beats volume.** The highest-leverage step: filter generated data
  for correctness. Executable domains (code, math) verify by running; open-ended
  domains use entailment checks, multi-sample consistency, or a second model as
  judge.
- **Rejection sampling / best-of-N.** Generate N candidates, keep the ones that
  pass verification. This converts compute into quality and is the workhorse of
  synthetic pipelines.
- **Distillation vs. self-improvement.** Distillation transfers from a stronger
  teacher; self-improvement (STaR-style) has the model generate its own rationales
  and keeps the ones leading to correct answers. Both need the same verification
  discipline.
- **Model collapse.** Repeated training on unfiltered model outputs narrows the
  distribution and amplifies errors. Mitigations: keep a substantial fraction of
  real data, filter aggressively, and regenerate from the strongest available
  teacher each round.
- **Contamination hygiene.** Synthetic eval-adjacent data must never leak into
  benchmarks. Track provenance: which teacher, which prompt, which filter version
  produced every example.
- **Teacher-student gap.** The student learns the teacher's verified outputs, not
  the teacher's capabilities. A student trained on traces can match the teacher's
  task performance without matching its generality — know which one you need.
- **Format diversity.** Varying output formats during generation prevents the
  student from overfitting to one teacher's stylistic template.

## Practical workflow

1. **Define the target capability precisely.** "Better reasoning" is not a spec.
   Write the task distribution: input types, difficulty mix, output format. The
   generation prompts derive from this spec.
2. **Seed for diversity.** Build a seed pool of topics, personas, difficulty
   levels, and edge cases. Sample seeds combinatorially so generation covers the
   space rather than clustering.
3. **Generate with a strong teacher.** Use the best model you can afford for
   generation — quality of the teacher bounds quality of the data. Generate
   multiple candidates per seed.
4. **Verify and filter.** Run domain-appropriate checks: execute code, check math
   with a symbolic or second-model verifier, use an LLM judge with a strict rubric
   for open-ended content. Typical keep rates: 20–60%. Log why examples were
   rejected.
5. **Deduplicate and balance.** Near-dup removal (embedding similarity), difficulty
   balancing, and format consistency. A synthetic dataset of 10k diverse, verified
   examples beats 100k repetitive ones.
6. **Mix with real data.** Blend synthetic with real examples (common starting
   ratios: 1:1 to 3:1 synthetic:real). Ablate the ratio — the optimal mix is
   task-specific.
7. **Train and measure the delta.** Fine-tune the student, evaluate on held-out
   benchmarks for the target capability AND on general benchmarks to catch
   regressions. If the delta is zero, fix the pipeline (usually verification or
   diversity), not the training hyperparameters.
8. **Record provenance.** Every example tagged with teacher model, prompt version,
   filter version. This is what makes the dataset auditable and regenerable.

Checklist for a synthetic dataset release:
- Verification method documented with measured precision on a human-labeled sample.
- Diversity measured (not asserted): embedding coverage, difficulty histogram.
- Real-data mix ratio chosen by ablation, not by default.
- Provenance recorded per example.
- Downstream delta measured on held-out evals, including regression checks.

## Common pitfalls

- **No verification.** The #1 failure. Unverified synthetic data teaches the
  student to imitate the teacher's mistakes confidently.
- **Teacher too weak.** A student can't exceed its teacher's verified quality. If
  the teacher can't solve the task reliably, its synthetic data is noise.
- **Diversity theater.** Generating 100k examples from 50 seeds gives you 100k
  near-duplicates. Seed combinatorics and explicit novelty pressure matter more
  than raw count.
- **Format overfitting.** The student learns the teacher's stylistic tics (headers,
  "Certainly!") rather than the substance. Vary output formats or strip
  boilerplate.
- **Eval contamination.** Synthetic data derived from or resembling benchmark items
  inflates scores without improving capability. Keep generation seeds disjoint
  from eval content.
- **Ignoring the real-data mix.** Pure synthetic training drifts. Blend with real
  data and ablate the ratio — don't guess it.
- **Weak verifier, strong claims.** A sloppy LLM judge passing bad examples
  poisons the dataset quietly. Measure your verifier's precision on human-labeled
  samples.
- **One-shot pipeline.** Building the pipeline once and never iterating. The first
  version's keep rate and diversity are always wrong — instrument, inspect
  rejects, improve.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →