**arXiv ID:** 2509.23982 **Authors:** Lucio La Cava, Andrea Tagarelli **Published:** 2025-09-28T17:16:16Z **Abstract:** Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Optimization typically require curated data and expensive optimization over billions of parameters, and eventually lead to persistent task-specific models. In th...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill toward-preferencealigned-large-language-models-via-residualbased-model-steering --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Toward Preferencealigned Large Language Models Via Residualbased Model Steering?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-toward-preferencealigned-large-language-models-via)More formats (shields.io, HTML) on the badges page.
# Toward Preference-aligned Large Language Models via Residual-based Model Steering
**arXiv ID:** 2509.23982
**Authors:** Lucio La Cava, Andrea Tagarelli
**Published:** 2025-09-28T17:16:16Z
**Abstract:**
Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Optimization typically require curated data and expensive optimization over billions of parameters, and eventually lead to persistent task-specific models. In this work, we introduce Preference alignment of Large Language Models via Residual Steering (PaLRS), a training-free method that exploits preference signals encoded in the residual streams of LLMs. From as few as one hundred preference pairs, PaLRS extracts lightweight, plug-and-play steering vectors that can be applied at inference time to push models toward preferred behaviors. We evaluate PaLRS on various small-to-medium-scale open-source LLMs, showing that PaLRS-aligned models achieve consistent gains on mathematical reasoning and code generation benchmarks while preserving baseline general-purpose performance. Moreover, when compared to models aligned with DPO and SimPO, they perform better with great time-savings. Our findings highlight that PaLRS offers an effective, much more efficient and flexible alternative to standard preference optimization pipelines, offering a training-free, plug-and-play mechanism for alignment with minimal data.
## Skill Description
This skill is generated from the arXiv paper: Toward Preference-aligned Large Language Models via Residual-based Model Steering (2509.23982).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2509.23982](http://arxiv.org/abs/2509.23982v2)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!