Toward Real-Time and Efficient Edge Intelligence: Advances and Challenges in Lightweight Machine Learning
DOI:
https://doi.org/10.61173/bwn0ez36Keywords:
Machine Learning, knowledge distillation, Lightweight modelAbstract
Deploying advanced Machine Learning (ML), particularly Deep Neural Networks (DNNs), on resource-constrained edge devices is crucial for realizing low-latency, privacy-preserving, and reliable edge intelligence applications. However, a significant gap exists between the high computational, memory, and energy demands of state-of-the-art models and the severe limitations inherent to edge hardware. This review systematically analyzes the field of lightweight ML for edge devices, aiming to bridge this gap. Methods: Focusing on the inference phase, the review critically examines three primary technical pillars: (1) Model Compression techniques, including knowledge distillation, network pruning (structured and unstructured), and quantization; (2) Efficient Neural Architecture Design of inherently compact models (e.g., MobileNet, ShuffleNet, EfficientNet series); and (3) Hardware-aware Optimization and Adaptation, encompassing operator fusion, dedicated inference engines, and leveraging heterogeneous systems. Results and Conclusion: The analysis highlights key achievements in reducing model size, complexity, and latency while maintaining accuracy. However, fundamental challenges persist, including the accuracy-efficiency tradeoff, hardware fragmentation, the memory wall bottleneck, and privacy/security concerns during deployment. Emerging solutions like neural-symbolic learning, adaptive federated learning, hardware-aware Neural Architecture Search (NAS), Processing-in-Memory (PIM) accelerators, and cross-stack co-design frameworks represent promising future directions. Overcoming these challenges is strategically vital for unlocking the full potential of ubiquitous, real-time edge intelligence.
References
[1] Wang Q, Jin G, Li Q, Wang K, Yang Z, Wang H. Industrial Edge Computing: Vision and Challenges. Information and Control, 2021, 50(3): 257-274.
[2] Hussain H, Tamizharasan P S, Rahul C S. Design possibilities and challenges of DNN models: a review on the perspective of end devices. Artificial Intelligence Review, 2022, 55: 5109–5167.
[3] Hadidi R, Cao J, Xie Y, Asgari B, Krishna T, Kim H. Characterizing the Deployment of Deep Neural Networks on Commercial Edge Devices. IEEE International Symposium on Workload Characterization, 2019: 35–48.
[4] Choudhary T, Mishra V, Goswami A, et al. A comprehensive survey on model compression and acceleration. Artificial Intelligence Review, 2020, 53: 5113–5155.
[5] Reed R. Pruning algorithms—a survey. IEEE Transactions on Neural Networks, 1993, 4(5): 740–747.
[6] Pham-Quoc C, Nguyen X Q, Thinh T N. Towards an FPGA- targeted Hardware/Software Co-design Framework for CNN- based Edge Computing. Mobile Networks and Applications, 2022, 27: 2024–2035.
[7] Li B. Software Framework for Embedded Neural Networks. In: Embedded Artificial Intelligence. Springer, 2024.
[8] Gou J, Yu B, Maybank S J, et al. Knowledge Distillation: A Survey. International Journal of Computer Vision, 2021, 129: 1789–1819.
[9] Yeom SK, Seegerer P, Lapuschkin S, Binder A, Wiedemann S, Müller KR, Samek W. Pruning by explaining: A novel criterion for deep neural network pruning. Pattern Recognition. 2021 Jul 1;115:107899.
[10] He Y, Xiao L. Structured Pruning for Deep Convolutional Dean&Francis ISSN 2959-6157 Neural Networks: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(5): 2900–2919.
[11] Anwar S, Hwang K, Sung W. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems (JETC). 2017 Feb 9;13(3):1-8.
[12] He Y, Kang G, Dong X, Fu Y, Yang Y. Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks. Proceedings of the 27th International Joint Conference on Artificial Intelligence, 2018: 2234–2240.
[13] Zhang X, Zhou X, Lin M, Sun J. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition 2018 (pp. 6848-6856).
[14] Liu X, Xu W, Wang Q, Zhang M. Energy-Efficient Computing Acceleration of Unmanned Aerial Vehicles Based on a CPU/FPGA/NPU Heterogeneous System. IEEE Internet of Things Journal, 2024, 11(16): 27126–27138.
[15] Tan T, Cao G. Deep Learning Video Analytics Through Edge Computing and Neural Processing Units on Mobile Devices. IEEE Transactions on Mobile Computing, 2023, 22(3): 1433–1448.
[16] Shuvo M M H, Islam S K, Cheng J, Morshed B I. Efficient Acceleration of Deep Learning Inference on Resource- Constrained Edge Devices: A Review. Proceedings of the IEEE, 2023, 111(1): 42–91.
[17] Yang P, Wang H, Yang J, Qian Z, Zhang Y, Lin X. Deep Learning Approaches for Similarity Computation: A Survey. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(12): 7893–7912.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
