**arXiv ID:** 2501.16337 **Authors:** Ivan Knunyants, Maryam Tavakol, Manolis Sifalakis, Yingfu Xu, Amirreza Yousefzadeh, Guangzhi Tang **Published:** 2025-01-09T19:13:03Z **Abstract:** The recent rise of Large Language Models (LLMs) has revolutionized the deep learning field. However, the desire to deploy LLMs on edge devices introduces energy efficiency and latency challenges. Recurrent LLM (R-LLM) architectures have proven effective in mitigating the quadratic complexity of self-attention,...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill explore-activation-sparsity-in-recurrent-llms-for-energyefficient-neuromorphic-computing --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Explore Activation Sparsity In Recurrent Llms For Energyefficient Neuromorphic Computing?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-explore-activation-sparsity-in-recurrent-llms-for)More formats (shields.io, HTML) on the badges page.
# Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing
**arXiv ID:** 2501.16337
**Authors:** Ivan Knunyants, Maryam Tavakol, Manolis Sifalakis, Yingfu Xu, Amirreza Yousefzadeh, Guangzhi Tang
**Published:** 2025-01-09T19:13:03Z
**Abstract:**
The recent rise of Large Language Models (LLMs) has revolutionized the deep learning field. However, the desire to deploy LLMs on edge devices introduces energy efficiency and latency challenges. Recurrent LLM (R-LLM) architectures have proven effective in mitigating the quadratic complexity of self-attention, making them a potential paradigm for computing on-edge neuromorphic processors. In this work, we propose a low-cost, training-free algorithm to sparsify R-LLMs' activations to enhance energy efficiency on neuromorphic hardware. Our approach capitalizes on the inherent structure of these models, rendering them well-suited for energy-constrained environments. Although primarily designed for R-LLMs, this method can be generalized to other LLM architectures, such as transformers, as demonstrated on the OPT model, achieving comparable sparsity and efficiency improvements. Empirical studies illustrate that our method significantly reduces computational demands while maintaining competitive accuracy across multiple zero-shot learning benchmarks. Additionally, hardware simulations with the SENECA neuromorphic processor underscore notable energy savings and latency improvements. These results pave the way for low-power, real-time neuromorphic deployment of LLMs and demonstrate the feasibility of training-free on-chip adaptation using activation sparsity.
## Skill Description
This skill is generated from the arXiv paper: Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing (2501.16337).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2501.16337](http://arxiv.org/abs/2501.16337v1)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!