Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Rlhf Pipeline

ASecurity

Build the full reinforcement-learning-from-human-feedback stack — preference data, reward model, and policy optimization.

2 stars
0 votes
0 copies
1 views
Added 9/29/2026
ai-agentsgoawstesting

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill rlhf-pipeline --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Rlhf Pipeline?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Rlhf Pipeline
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-rlhf-pipeline/badge)](https://www.skillsdirectory.com/skills/aicodedecode-rlhf-pipeline)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: rlhf-pipeline
description: Build the full reinforcement-learning-from-human-feedback stack — preference data, reward model, and policy optimization.
category: ai-research
---

## Overview

RLHF aligns a language model to human preferences in three stages. First, collect
preference data: human annotators rank or choose between pairs of model responses
to prompts. Second, train a reward model on these comparisons, typically with a
Bradley-Terry loss that predicts which response a human would prefer. Third,
optimize the policy — usually with PPO (proximal policy optimization) — to maximize
reward while staying close to the original model via a KL penalty, preventing
collapse into reward-hacking gibberish.

Variants replace pieces of this stack: DPO skips the reward model and optimizes
the policy directly on preference pairs; RLAIF uses AI-generated preferences;
online RLHF re-collects preferences as the policy changes. The classic three-stage
pipeline remains the reference design because each stage is independently
inspectable: you can audit the preference data, probe the reward model, and
monitor the KL budget during policy training.

The economic reality: RLHF is the most expensive common alignment method, in
labeling cost and engineering complexity. It earns that cost when nuanced human
judgment — tone, helpfulness trade-offs, refusal style — is the actual product
requirement.

## When to use

- Aligning a base or instruction-tuned model to nuanced human preferences that can't
  be captured by supervised data alone (tone, helpfulness trade-offs, refusal
  style).
- Building an assistant product where "which response is better" judgments matter
  more than "which is correct."
- Reducing harmful or unwanted outputs while keeping the model helpful, via a
  reward model that encodes those trade-offs.
- Researching alignment: each stage is a separate object of study (labeler
  agreement, reward model generalization, optimization dynamics).
- When you have the labeling budget and infrastructure for a multi-stage pipeline
  and the scale to justify it.
- As the final polish stage after SFT, when supervised data has taken you as far
  as demonstrations can.

## Core concepts

- **Preference pairs**: (prompt, chosen, rejected) triples. Quality dominates
  quantity: clear, well-annotated comparisons beat large volumes of noisy ones.
  Annotator guidelines are the real specification of your model's values.
- **Reward model**: usually the policy-sized (or smaller) model with a scalar head,
  trained to score responses. Evaluated by accuracy on held-out human preferences
  and by calibration on out-of-distribution prompts.
- **Bradley-Terry model**: the standard loss — the probability that response A is
  preferred over B is a sigmoid of their score difference. Simple, but assumes
  transitive, noise-free preferences; real data isn't.
- **PPO with KL penalty**: the policy update maximizes expected reward minus
  β·KL(policy || reference). The KL term is the guardrail against reward hacking.
  Track reward and KL jointly — rising reward with rising KL is a warning sign,
  not success.
- **Reward hacking**: the policy learns to exploit quirks of the reward model
  (verbosity, sycophancy, confident tone) rather than genuine quality. The single
  biggest failure mode of the whole pipeline.
- **Online vs. offline**: offline RLHF trains on a fixed preference dataset;
  online variants (iterative RLHF) generate new responses from the current policy
  and get fresh preferences, correcting distribution shift.
- **KL budget**: the total allowed divergence from the reference, spent across
  training. Spend it deliberately — monitor cumulative KL, not just the
  coefficient.
- **Annotator agreement**: the noise floor of your signal. Low inter-annotator
  agreement means the reward model learns noise; fix guidelines before scaling
  collection.

## Practical workflow

1. **Write annotation guidelines.** Specify what "better" means for your use case:
   factuality first? refusal style? tone? Include edge-case examples. Pilot with
   200–500 examples and measure inter-annotator agreement before scaling.
2. **Collect preference data.** Sample prompts covering the target distribution;
   generate 2–4 responses per prompt from the current policy; have annotators rank
   or pick best/worst. Aim for tens of thousands of pairs minimum for a serious
   run.
3. **Train the reward model.** Fine-tune from the policy or a smaller model with a
   scalar head. Validate on held-out pairs and on adversarial prompts; check that
   scores correlate with human judgment on a fresh sample.
4. **Probe the reward model.** Before PPO, test it on edge cases: does it prefer
   longer answers regardless of quality? Does it penalize refusals correctly? Fix
   what you find — PPO will amplify every flaw.
5. **Run PPO.** Standard hyperparameters: KL coefficient β tuned so KL stays in a
   target band, learning rate lower than SFT, batch sizes large enough for stable
   advantage estimates. Log reward, KL, response length, and win-rate against the
   base policy throughout.
6. **Evaluate holistically.** Reward score is not the metric — human win-rate on a
   blind eval set, plus targeted safety evals, plus capability regression checks
   (the aligned model shouldn't forget facts or reasoning).
7. **Iterate online if needed.** If you see distribution shift (the policy exploits
   regions the reward model never saw), collect fresh preferences on current-policy
   outputs and repeat.

Checklist before shipping an RLHF model:
- Held-out human win-rate beats the SFT baseline by a meaningful margin.
- Capability benchmarks haven't regressed (no "alignment tax" surprises).
- Reward hacking probes: verbosity, sycophancy, refusal-bait prompts tested.
- KL divergence stayed within the planned budget through training.
- Annotator guidelines versioned and agreement metrics recorded.

## Common pitfalls

- **Optimizing the reward, not the goal.** Goodhart's law is the default outcome.
  Treat the reward model as a proxy and validate with humans continuously.
- **Bad preference data.** Inconsistent annotators, unclear guidelines, or prompts
  that don't match deployment produce a reward model that encodes the wrong values
  precisely.
- **KL too loose or too tight.** Too loose → reward hacking and degenerate outputs.
  Too tight → the RL stage changes nothing and you've paid for nothing.
- **Ignoring length bias.** Reward models systematically prefer longer responses.
  Control for length in evaluation or the policy learns to ramble.
- **Single-turn optimization.** PPO on single responses can degrade multi-turn
  behavior (e.g., the model becomes agreeable rather than truthful across a
  conversation). Evaluate multi-turn explicitly.
- **Skipping the SFT stage.** RLHF on a raw base model is unstable; a solid
  supervised fine-tune first gives the policy a sensible starting distribution.
- **Reward model never probed.** Going straight from reward training to PPO without
  adversarially testing the reward model. Its flaws become the policy's features.
- **No online correction.** Running one offline round and declaring victory while
  the policy has drifted off the reward model's training distribution.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →