Skip to content
Back to skills

Nlp Alignment

ASecurity

Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety.

  • 10 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgo

Works with

  • cli

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add FOURTEEN1416/academic-agent-toolkit --skill nlp-alignment --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Nlp Alignment?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Nlp Alignment
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fourteen1416-nlp-alignment/badge)](https://www.skillsdirectory.com/skills/fourteen1416-nlp-alignment)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nlp-alignment
description: Best practices for LLM alignment techniques including RLHF, DPO, and instruction tuning. Use when working on alignment or safety.
metadata:
  category: domain
  trigger-keywords: "alignment,rlhf,dpo,reward model,preference,instruction tuning,safety"
  applicable-stages: "9,10"
  priority: "4"
  version: "1.0"
  author: researchclaw
  references: "Ouyang et al., Training language models to follow instructions, NeurIPS 2022; Rafailov et al., DPO, NeurIPS 2023"
---

## LLM Alignment Best Practice
Methods:
- RLHF: Train reward model → PPO fine-tuning (complex but powerful)
- DPO: Direct preference optimization (simpler, no reward model needed)
- GRPO: Group relative policy optimization
- SFT: Supervised fine-tuning as alignment baseline

Training recipe:
- Start with SFT on high-quality instruction data
- DPO: lr=5e-7, beta=0.1, batch_size=64
- PPO: lr=1e-6, clip=0.2, KL coeff=0.02
- Use reference model for KL penalty
- Evaluate on safety benchmarks (TruthfulQA, BBQ, etc.)

Common pitfalls:
- Reward hacking: model finds shortcuts to high reward
- Mode collapse: model generates repetitive outputs
- Catastrophic forgetting: loses general capabilities

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…