**arXiv ID:** 2501.15014 **Authors:** Jacob Sander, Achraf Cohen, Venkat R. Dasari, Brent Venable, Brian Jalaian **Published:** 2025-01-25T01:37:03Z **Abstract:** Resource-constrained edge deployments demand AI solutions that balance high performance with stringent compute, memory, and energy limitations. In this survey, we present a comprehensive overview of the primary strategies for accelerating deep learning models under such constraints. First, we examine model compression techniques-pru...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill on-accelerating-edge-ai-optimizing-resourceconstrained-environments --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of On Accelerating Edge Ai Optimizing Resourceconstrained Environments?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-on-accelerating-edge-ai-optimizing-resourceconstra)More formats (shields.io, HTML) on the badges page.
# On Accelerating Edge AI: Optimizing Resource-Constrained Environments
**arXiv ID:** 2501.15014
**Authors:** Jacob Sander, Achraf Cohen, Venkat R. Dasari, Brent Venable, Brian Jalaian
**Published:** 2025-01-25T01:37:03Z
**Abstract:**
Resource-constrained edge deployments demand AI solutions that balance high performance with stringent compute, memory, and energy limitations. In this survey, we present a comprehensive overview of the primary strategies for accelerating deep learning models under such constraints. First, we examine model compression techniques-pruning, quantization, tensor decomposition, and knowledge distillation-that streamline large models into smaller, faster, and more efficient variants. Next, we explore Neural Architecture Search (NAS), a class of automated methods that discover architectures inherently optimized for particular tasks and hardware budgets. We then discuss compiler and deployment frameworks, such as TVM, TensorRT, and OpenVINO, which provide hardware-tailored optimizations at inference time. By integrating these three pillars into unified pipelines, practitioners can achieve multi-objective goals, including latency reduction, memory savings, and energy efficiency-all while maintaining competitive accuracy. We also highlight emerging frontiers in hierarchical NAS, neurosymbolic approaches, and advanced distillation tailored to large language models, underscoring open challenges like pre-training pruning for massive networks. Our survey offers practical insights, identifies current research gaps, and outlines promising directions for building scalable, platform-independent frameworks to accelerate deep learning models at the edge.
## Skill Description
This skill is generated from the arXiv paper: On Accelerating Edge AI: Optimizing Resource-Constrained Environments (2501.15014).
## How to Use
[To be filled in by the user or by future automation]
## References
- [arXiv:2501.15014](http://arxiv.org/abs/2501.15014v2)
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!