**arXiv ID:** 1509.03005 **Authors:** David Balduzzi, Muhammad Ghifary **Published:** 2015-09-10T04:14:54Z **Abstract:** This paper proposes GProp, a deep reinforcement learning algorithm for continuous policies with compatible function approximation. The algorithm is based on two innovations. Firstly, we present a temporal-difference based method for learning the gradient of the value-function. Secondly, we present the deviator-actor-critic (DAC) model, which comprises three neural networks ...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill compatible-value-gradients-for-reinforcement-learning-of-continuous-deep-policies --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Compatible Value Gradients For Reinforcement Learning Of Continuous Deep Policies?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-compatible-value-gradients-for-reinforcement-learn)More formats (shields.io, HTML) on the badges page.
# Compatible Value Gradients for Reinforcement Learning of Continuous Deep Policies
**arXiv ID:** 1509.03005
**Authors:** David Balduzzi, Muhammad Ghifary
**Published:** 2015-09-10T04:14:54Z
**Abstract:**
This paper proposes GProp, a deep reinforcement learning algorithm for continuous policies with compatible function approximation. The algorithm is based on two innovations. Firstly, we present a temporal-difference based method for learning the gradient of the value-function. Secondly, we present the deviator-actor-critic (DAC) model, which comprises three neural networks that estimate the value function, its gradient, and determine the actor's policy respectively. We evaluate GProp on two challenging tasks: a contextual bandit problem constructed from nonparametric regression datasets that is designed to probe the ability of reinforcement learning algorithms to accurately estimate gradients; and the octopus arm, a challenging reinforcement learning benchmark. GProp is competitive with fully supervised methods on the bandit task and achieves the best performance to date on the octopus arm.
## Skill Description
This skill is generated from the arXiv paper: Compatible Value Gradients for Reinforcement Learning of Continuous Deep Policies (1509.03005).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:1509.03005](http://arxiv.org/abs/1509.03005v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!