**arXiv ID:** 1705.10993 **Authors:** Julien Perez, Tomi Silander **Published:** 2017-05-31T09:00:44Z **Abstract:** Partially observable environments present an important open challenge in the domain of sequential control learning with delayed rewards. Despite numerous attempts during the two last decades, the majority of reinforcement learning algorithms and associated approximate models, applied to this context, still assume Markovian state transitions. In this paper, we explore the use of ...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill nonmarkovian-control-with-gated-endtoend-memory-policy-networks --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Nonmarkovian Control With Gated Endtoend Memory Policy Networks?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-nonmarkovian-control-with-gated-endtoend-memory-po)More formats (shields.io, HTML) on the badges page.
# Non-Markovian Control with Gated End-to-End Memory Policy Networks
**arXiv ID:** 1705.10993
**Authors:** Julien Perez, Tomi Silander
**Published:** 2017-05-31T09:00:44Z
**Abstract:**
Partially observable environments present an important open challenge in the domain of sequential control learning with delayed rewards. Despite numerous attempts during the two last decades, the majority of reinforcement learning algorithms and associated approximate models, applied to this context, still assume Markovian state transitions. In this paper, we explore the use of a recently proposed attention-based model, the Gated End-to-End Memory Network, for sequential control. We call the resulting model the Gated End-to-End Memory Policy Network. More precisely, we use a model-free value-based algorithm to learn policies for partially observed domains using this memory-enhanced neural network. This model is end-to-end learnable and it features unbounded memory. Indeed, because of its attention mechanism and associated non-parametric memory, the proposed model allows us to define an attention mechanism over the observation stream unlike recurrent models. We show encouraging results that illustrate the capability of our attention-based model in the context of the continuous-state non-stationary control problem of stock trading. We also present an OpenAI Gym environment for simulated stock exchange and explain its relevance as a benchmark for the field of non-Markovian decision process learning.
## Skill Description
This skill is generated from the arXiv paper: Non-Markovian Control with Gated End-to-End Memory Policy Networks (1705.10993).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:1705.10993](http://arxiv.org/abs/1705.10993v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!