**arXiv ID:** 2601.21503 **Authors:** Junhong Cai, Guiqin Wang, Kejie Zhao, Jianxiong Tang, Xiang Wang, Luziwei Leng, Ran Cheng, Yuxin Ma, Qinghai Guo **Published:** 2026-01-29T10:21:28Z **Abstract:** Large Language Models (LLMs) excel across diverse domains but suffer from high energy costs due to quadratic attention and dense Feed-Forward Network (FFN) operations. To address these issues, we propose Module-aware Architecture Refinement (MAR), a two-stage framework that integrates State Spac...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill mar-efficient-large-language-models-via-moduleaware-architecture-refinement --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mar Efficient Large Language Models Via Moduleaware Architecture Refinement?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-mar-efficient-large-language-models-via-moduleawar)More formats (shields.io, HTML) on the badges page.
# MAR: Efficient Large Language Models via Module-aware Architecture Refinement
**arXiv ID:** 2601.21503
**Authors:** Junhong Cai, Guiqin Wang, Kejie Zhao, Jianxiong Tang, Xiang Wang, Luziwei Leng, Ran Cheng, Yuxin Ma, Qinghai Guo
**Published:** 2026-01-29T10:21:28Z
**Abstract:**
Large Language Models (LLMs) excel across diverse domains but suffer from high energy costs due to quadratic attention and dense Feed-Forward Network (FFN) operations. To address these issues, we propose Module-aware Architecture Refinement (MAR), a two-stage framework that integrates State Space Models (SSMs) for linear-time sequence modeling and applies activation sparsification to reduce FFN costs. In addition, to mitigate low information density and temporal mismatch in integrating Spiking Neural Networks (SNNs) with SSMs, we design the Adaptive Ternary Multi-step Neuron (ATMN) and the Spike-aware Bidirectional Distillation Strategy (SBDS). Extensive experiments demonstrate that MAR effectively restores the performance of its dense counterpart under constrained resources while substantially reducing inference energy consumption. Furthermore, it outperforms efficient models of comparable or even larger scale, underscoring its potential for building efficient and practical LLMs.
## Skill Description
This skill is generated from the arXiv paper: MAR: Efficient Large Language Models via Module-aware Architecture Refinement (2601.21503).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2601.21503](http://arxiv.org/abs/2601.21503v1)
Is this your skill, or is something wrong with this listing? . Author removals are honored within 72 hours.
No comments yet. Be the first to comment!