**arXiv ID:** 2304.08488 **Authors:** Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, Deepak Pathak **Published:** 2023-04-17T17:59:34Z **Abstract:** Building a robot that can understand and learn to interact by watching humans has inspired several vision problems. However, despite some successful results on static datasets, it remains unclear how current models can be used on a robot directly. In this paper, we aim to bridge this gap by leveraging videos of human interactions in an en...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill affordances-from-human-videos-as-a-versatile-representation-for-robotics --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Affordances From Human Videos As A Versatile Representation For Robotics?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-affordances-from-human-videos-as-a-versatile-repre)More formats (shields.io, HTML) on the badges page.
# Affordances from Human Videos as a Versatile Representation for Robotics
**arXiv ID:** 2304.08488
**Authors:** Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, Deepak Pathak
**Published:** 2023-04-17T17:59:34Z
**Abstract:**
Building a robot that can understand and learn to interact by watching humans has inspired several vision problems. However, despite some successful results on static datasets, it remains unclear how current models can be used on a robot directly. In this paper, we aim to bridge this gap by leveraging videos of human interactions in an environment centric manner. Utilizing internet videos of human behavior, we train a visual affordance model that estimates where and how in the scene a human is likely to interact. The structure of these behavioral affordances directly enables the robot to perform many complex tasks. We show how to seamlessly integrate our affordance model with four robot learning paradigms including offline imitation learning, exploration, goal-conditioned learning, and action parameterization for reinforcement learning. We show the efficacy of our approach, which we call VRB, across 4 real world environments, over 10 different tasks, and 2 robotic platforms operating in the wild. Results, visualizations and videos at https://robo-affordances.github.io/
## Skill Description
This skill is generated from the arXiv paper: Affordances from Human Videos as a Versatile Representation for Robotics (2304.08488).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2304.08488](http://arxiv.org/abs/2304.08488v1)
Is this your skill, or is something wrong with this listing? . Author removals are honored within 72 hours.
No comments yet. Be the first to comment!