**arXiv ID:** 2204.03655 **Authors:** Bryan Lim, Alexander Reichenbach, Antoine Cully **Published:** 2022-04-07T14:07:51Z **Abstract:** Quality-Diversity (QD) algorithms can discover large and complex behavioural repertoires consisting of both diverse and high-performing skills. However, the generation of behavioural repertoires has mainly been limited to simulation environments instead of real-world learning. This is because existing QD algorithms need large numbers of evaluations as well as...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill learning-to-walk-autonomously-via-resetfree-qualitydiversity --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Learning To Walk Autonomously Via Resetfree Qualitydiversity?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-learning-to-walk-autonomously-via-resetfree-qualit)More formats (shields.io, HTML) on the badges page.
# Learning to Walk Autonomously via Reset-Free Quality-Diversity
**arXiv ID:** 2204.03655
**Authors:** Bryan Lim, Alexander Reichenbach, Antoine Cully
**Published:** 2022-04-07T14:07:51Z
**Abstract:**
Quality-Diversity (QD) algorithms can discover large and complex behavioural repertoires consisting of both diverse and high-performing skills. However, the generation of behavioural repertoires has mainly been limited to simulation environments instead of real-world learning. This is because existing QD algorithms need large numbers of evaluations as well as episodic resets, which require manual human supervision and interventions. This paper proposes Reset-Free Quality-Diversity optimization (RF-QD) as a step towards autonomous learning for robotics in open-ended environments. We build on Dynamics-Aware Quality-Diversity (DA-QD) and introduce a behaviour selection policy that leverages the diversity of the imagined repertoire and environmental information to intelligently select of behaviours that can act as automatic resets. We demonstrate this through a task of learning to walk within defined training zones with obstacles. Our experiments show that we can learn full repertoires of legged locomotion controllers autonomously without manual resets with high sample efficiency in spite of harsh safety constraints. Finally, using an ablation of different target objectives, we show that it is important for RF-QD to have diverse types solutions available for the behaviour selection policy over solutions optimised with a specific objective. Videos and code available at https://sites.google.com/view/rf-qd.
## Skill Description
This skill is generated from the arXiv paper: Learning to Walk Autonomously via Reset-Free Quality-Diversity (2204.03655).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2204.03655](http://arxiv.org/abs/2204.03655v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!