**arXiv ID:** 2208.01191 **Authors:** Yunfan Zhao, Qingkai Pan, Krzysztof Choromanski, Deepali Jain, Vikas Sindhwani **Published:** 2022-08-02T01:23:50Z **Abstract:** We present a new class of structured reinforcement learning policy-architectures, Implicit Two-Tower (ITT) policies, where the actions are chosen based on the attention scores of their learnable latent representations with those of the input states. By explicitly disentangling action from state processing in the policy stack, we...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill implicit-twotower-policies --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Implicit Twotower Policies?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-implicit-twotower-policies)More formats (shields.io, HTML) on the badges page.
# Implicit Two-Tower Policies
**arXiv ID:** 2208.01191
**Authors:** Yunfan Zhao, Qingkai Pan, Krzysztof Choromanski, Deepali Jain, Vikas Sindhwani
**Published:** 2022-08-02T01:23:50Z
**Abstract:**
We present a new class of structured reinforcement learning policy-architectures, Implicit Two-Tower (ITT) policies, where the actions are chosen based on the attention scores of their learnable latent representations with those of the input states. By explicitly disentangling action from state processing in the policy stack, we achieve two main goals: substantial computational gains and better performance. Our architectures are compatible with both: discrete and continuous action spaces. By conducting tests on 15 environments from OpenAI Gym and DeepMind Control Suite, we show that ITT-architectures are particularly suited for blackbox/evolutionary optimization and the corresponding policy training algorithms outperform their vanilla unstructured implicit counterparts as well as commonly used explicit policies. We complement our analysis by showing how techniques such as hashing and lazy tower updates, critically relying on the two-tower structure of ITTs, can be applied to obtain additional computational improvements.
## Skill Description
This skill is generated from the arXiv paper: Implicit Two-Tower Policies (2208.01191).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2208.01191](http://arxiv.org/abs/2208.01191v2)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!