**arXiv ID:** 1811.06521 **Authors:** Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei **Published:** 2018-11-15T18:33:43Z **Abstract:** To solve complex real-world problems with reinforcement learning, we cannot rely on manually specified reward functions. Instead, we can have humans communicate an objective to the agent directly. In this work, we combine two approaches to learning from human feedback: expert demonstrations and trajectory preferences. We train...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill reward-learning-from-human-preferences-and-demonstrations-in-atari --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Reward Learning From Human Preferences And Demonstrations In Atari?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-reward-learning-from-human-preferences-and-demonst)More formats (shields.io, HTML) on the badges page.
# Reward learning from human preferences and demonstrations in Atari
**arXiv ID:** 1811.06521
**Authors:** Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei
**Published:** 2018-11-15T18:33:43Z
**Abstract:**
To solve complex real-world problems with reinforcement learning, we cannot rely on manually specified reward functions. Instead, we can have humans communicate an objective to the agent directly. In this work, we combine two approaches to learning from human feedback: expert demonstrations and trajectory preferences. We train a deep neural network to model the reward function and use its predicted reward to train an DQN-based deep reinforcement learning agent on 9 Atari games. Our approach beats the imitation learning baseline in 7 games and achieves strictly superhuman performance on 2 games without using game rewards. Additionally, we investigate the goodness of fit of the reward model, present some reward hacking problems, and study the effects of noise in the human labels.
## Skill Description
This skill is generated from the arXiv paper: Reward learning from human preferences and demonstrations in Atari (1811.06521).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:1811.06521](http://arxiv.org/abs/1811.06521v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!