Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.
Scanned 9/29/2026
npx -y skills add aicodedecode/awesome-muse-skills --skill synthetic-data --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Synthetic Data?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/aicodedecode-synthetic-data)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: synthetic-data
description: Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.
category: ai-research
---
## Overview
Synthetic data is training data generated by models rather than collected from
humans or the world: instruction-response pairs, reasoning traces, code with
tests, preference judgments, tool-use trajectories. It's the engine behind most
modern post-training — distillation, self-improvement loops, and domain adaptation
all run on synthetic data. Done well, it converts a strong teacher model's
capabilities into targeted training signal for a student. Done poorly, it produces
fluent garbage that teaches the student the teacher's failure modes.
The core discipline is treating synthetic generation as a data pipeline with
quality gates, not as "prompt a big model and hope." Every synthetic dataset needs:
a generation step with controlled diversity, a verification step that filters for
correctness (execution, entailment checks, cross-model agreement), and a validation
step measuring downstream effect on the student. The failure mode to fear is model
collapse — training on unfiltered synthetic data degrades quality over
generations — which is why verification and mixing with real data matter.
Synthetic data is leverage: it multiplies a strong teacher into arbitrary
quantities of targeted training signal. Like all leverage, it amplifies errors
too. The pipeline's job is to keep the signal while discarding the errors.
## When to use
- Post-training a model for a domain where human data is scarce or expensive
(specialized reasoning, low-resource languages, niche tools).
- Distilling capabilities from a stronger teacher into a smaller, cheaper student.
- Generating instruction-tuning data at a scale humans can't match.
- Creating adversarial or edge-case examples to patch specific weaknesses found in
evals.
- Bootstrapping preference data (pairs judged by a strong model) for DPO or
reward modeling.
- Augmenting thin slices of real data (rare classes, edge cases) where collection
is impractical.
## Core concepts
- **Generation with diversity control.** Naive sampling collapses to the teacher's
modal outputs. Use temperature variation, persona/topic seeding, and explicit
diversity prompts ("generate a problem unlike these examples") to cover the
space.
- **Verification beats volume.** The highest-leverage step: filter generated data
for correctness. Executable domains (code, math) verify by running; open-ended
domains use entailment checks, multi-sample consistency, or a second model as
judge.
- **Rejection sampling / best-of-N.** Generate N candidates, keep the ones that
pass verification. This converts compute into quality and is the workhorse of
synthetic pipelines.
- **Distillation vs. self-improvement.** Distillation transfers from a stronger
teacher; self-improvement (STaR-style) has the model generate its own rationales
and keeps the ones leading to correct answers. Both need the same verification
discipline.
- **Model collapse.** Repeated training on unfiltered model outputs narrows the
distribution and amplifies errors. Mitigations: keep a substantial fraction of
real data, filter aggressively, and regenerate from the strongest available
teacher each round.
- **Contamination hygiene.** Synthetic eval-adjacent data must never leak into
benchmarks. Track provenance: which teacher, which prompt, which filter version
produced every example.
- **Teacher-student gap.** The student learns the teacher's verified outputs, not
the teacher's capabilities. A student trained on traces can match the teacher's
task performance without matching its generality — know which one you need.
- **Format diversity.** Varying output formats during generation prevents the
student from overfitting to one teacher's stylistic template.
## Practical workflow
1. **Define the target capability precisely.** "Better reasoning" is not a spec.
Write the task distribution: input types, difficulty mix, output format. The
generation prompts derive from this spec.
2. **Seed for diversity.** Build a seed pool of topics, personas, difficulty
levels, and edge cases. Sample seeds combinatorially so generation covers the
space rather than clustering.
3. **Generate with a strong teacher.** Use the best model you can afford for
generation — quality of the teacher bounds quality of the data. Generate
multiple candidates per seed.
4. **Verify and filter.** Run domain-appropriate checks: execute code, check math
with a symbolic or second-model verifier, use an LLM judge with a strict rubric
for open-ended content. Typical keep rates: 20–60%. Log why examples were
rejected.
5. **Deduplicate and balance.** Near-dup removal (embedding similarity), difficulty
balancing, and format consistency. A synthetic dataset of 10k diverse, verified
examples beats 100k repetitive ones.
6. **Mix with real data.** Blend synthetic with real examples (common starting
ratios: 1:1 to 3:1 synthetic:real). Ablate the ratio — the optimal mix is
task-specific.
7. **Train and measure the delta.** Fine-tune the student, evaluate on held-out
benchmarks for the target capability AND on general benchmarks to catch
regressions. If the delta is zero, fix the pipeline (usually verification or
diversity), not the training hyperparameters.
8. **Record provenance.** Every example tagged with teacher model, prompt version,
filter version. This is what makes the dataset auditable and regenerable.
Checklist for a synthetic dataset release:
- Verification method documented with measured precision on a human-labeled sample.
- Diversity measured (not asserted): embedding coverage, difficulty histogram.
- Real-data mix ratio chosen by ablation, not by default.
- Provenance recorded per example.
- Downstream delta measured on held-out evals, including regression checks.
## Common pitfalls
- **No verification.** The #1 failure. Unverified synthetic data teaches the
student to imitate the teacher's mistakes confidently.
- **Teacher too weak.** A student can't exceed its teacher's verified quality. If
the teacher can't solve the task reliably, its synthetic data is noise.
- **Diversity theater.** Generating 100k examples from 50 seeds gives you 100k
near-duplicates. Seed combinatorics and explicit novelty pressure matter more
than raw count.
- **Format overfitting.** The student learns the teacher's stylistic tics (headers,
"Certainly!") rather than the substance. Vary output formats or strip
boilerplate.
- **Eval contamination.** Synthetic data derived from or resembling benchmark items
inflates scores without improving capability. Keep generation seeds disjoint
from eval content.
- **Ignoring the real-data mix.** Pure synthetic training drifts. Blend with real
data and ablate the ratio — don't guess it.
- **Weak verifier, strong claims.** A sloppy LLM judge passing bad examples
poisons the dataset quietly. Measure your verifier's precision on human-labeled
samples.
- **One-shot pipeline.** Building the pipeline once and never iterating. The first
version's keep rate and diversity are always wrong — instrument, inspect
rejects, improve.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!