**arXiv ID:** 2101.03958 **Authors:** John D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real, Sergey Levine, Quoc V. Le, Honglak Lee, Aleksandra Faust **Published:** 2021-01-08T18:55:07Z **Abstract:** We propose a method for meta-learning reinforcement learning algorithms by searching over the space of computational graphs which compute the loss function for a value-based model-free RL agent to optimize. The learned algorithms are domain-agnostic and can generalize to new environments not s...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill evolving-reinforcement-learning-algorithms --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Evolving Reinforcement Learning Algorithms?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-evolving-reinforcement-learning-algorithms)More formats (shields.io, HTML) on the badges page.
# Evolving Reinforcement Learning Algorithms
**arXiv ID:** 2101.03958
**Authors:** John D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real, Sergey Levine, Quoc V. Le, Honglak Lee, Aleksandra Faust
**Published:** 2021-01-08T18:55:07Z
**Abstract:**
We propose a method for meta-learning reinforcement learning algorithms by searching over the space of computational graphs which compute the loss function for a value-based model-free RL agent to optimize. The learned algorithms are domain-agnostic and can generalize to new environments not seen during training. Our method can both learn from scratch and bootstrap off known existing algorithms, like DQN, enabling interpretable modifications which improve performance. Learning from scratch on simple classical control and gridworld tasks, our method rediscovers the temporal-difference (TD) algorithm. Bootstrapped from DQN, we highlight two learned algorithms which obtain good generalization performance over other classical control tasks, gridworld type tasks, and Atari games. The analysis of the learned algorithm behavior shows resemblance to recently proposed RL algorithms that address overestimation in value-based methods.
## Skill Description
This skill is generated from the arXiv paper: Evolving Reinforcement Learning Algorithms (2101.03958).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2101.03958](http://arxiv.org/abs/2101.03958v6)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!