**arXiv ID:** 1506.01911 **Authors:** Lionel Pigou, Aäron van den Oord, Sander Dieleman, Mieke Van Herreweghe, Joni Dambre **Published:** 2015-06-05T13:43:01Z **Abstract:** Recent studies have demonstrated the power of recurrent neural networks for machine translation, image captioning and speech recognition. For the task of capturing temporal structure in video, however, there still remain numerous open research questions. Current research suggests using a simple temporal feature pooling str...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill beyond-temporal-pooling-recurrence-and-temporal-convolutions-for-gesture-recognition-in-video --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Beyond Temporal Pooling Recurrence And Temporal Convolutions For Gesture Recognition In Video?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-beyond-temporal-pooling-recurrence-and-temporal-co)More formats (shields.io, HTML) on the badges page.
# Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video
**arXiv ID:** 1506.01911
**Authors:** Lionel Pigou, Aäron van den Oord, Sander Dieleman, Mieke Van Herreweghe, Joni Dambre
**Published:** 2015-06-05T13:43:01Z
**Abstract:**
Recent studies have demonstrated the power of recurrent neural networks for machine translation, image captioning and speech recognition. For the task of capturing temporal structure in video, however, there still remain numerous open research questions. Current research suggests using a simple temporal feature pooling strategy to take into account the temporal aspect of video. We demonstrate that this method is not sufficient for gesture recognition, where temporal information is more discriminative compared to general video classification tasks. We explore deep architectures for gesture recognition in video and propose a new end-to-end trainable neural network architecture incorporating temporal convolutions and bidirectional recurrence. Our main contributions are twofold; first, we show that recurrence is crucial for this task; second, we show that adding temporal convolutions leads to significant improvements. We evaluate the different approaches on the Montalbano gesture recognition dataset, where we achieve state-of-the-art results.
## Skill Description
This skill is generated from the arXiv paper: Beyond Temporal Pooling: Recurrence and Temporal Convolutions for Gesture Recognition in Video (1506.01911).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:1506.01911](http://arxiv.org/abs/1506.01911v3)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!