**arXiv ID:** 1812.00045 **Authors:** Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor **Published:** 2018-11-30T20:37:17Z **Abstract:** Deep reinforcement learning (DRL) has achieved great successes in recent years with the help of novel methods and higher compute power. However, there are still several challenges to be addressed such as convergence to locally optimal policies and long training times. In this paper, firstly, we augment Asynchronous Advantage Actor-Critic (A3C) method wi...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill using-monte-carlo-tree-search-as-a-demonstrator-within-asynchronous-deep-rl --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Using Monte Carlo Tree Search As A Demonstrator Within Asynchronous Deep Rl?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-using-monte-carlo-tree-search-as-a-demonstrator-wi)More formats (shields.io, HTML) on the badges page.
# Using Monte Carlo Tree Search as a Demonstrator within Asynchronous Deep RL
**arXiv ID:** 1812.00045
**Authors:** Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor
**Published:** 2018-11-30T20:37:17Z
**Abstract:**
Deep reinforcement learning (DRL) has achieved great successes in recent years with the help of novel methods and higher compute power. However, there are still several challenges to be addressed such as convergence to locally optimal policies and long training times. In this paper, firstly, we augment Asynchronous Advantage Actor-Critic (A3C) method with a novel self-supervised auxiliary task, i.e. \emph{Terminal Prediction}, measuring temporal closeness to terminal states, namely A3C-TP. Secondly, we propose a new framework where planning algorithms such as Monte Carlo tree search or other sources of (simulated) demonstrators can be integrated to asynchronous distributed DRL methods. Compared to vanilla A3C, our proposed methods both learn faster and converge to better policies on a two-player mini version of the Pommerman game.
## Skill Description
This skill is generated from the arXiv paper: Using Monte Carlo Tree Search as a Demonstrator within Asynchronous Deep RL (1812.00045).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:1812.00045](http://arxiv.org/abs/1812.00045v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!