Toward Real-Time and Efficient Edge Intelligence: Advances and Challenges in Lightweight Machine Learning

Authors

  • Xiang Gao

DOI:

https://doi.org/10.61173/bwn0ez36

Keywords:

Machine Learning, knowledge distillation, Lightweight model

Abstract

Deploying advanced Machine Learning (ML), particularly Deep Neural Networks (DNNs), on resource-constrained edge devices is crucial for realizing low-latency, privacy-preserving, and reliable edge intelligence applications. However, a significant gap exists between the high computational, memory, and energy demands of state-of-the-art models and the severe limitations inherent to edge hardware. This review systematically analyzes the field of lightweight ML for edge devices, aiming to bridge this gap. Methods: Focusing on the inference phase, the review critically examines three primary technical pillars: (1) Model Compression techniques, including knowledge distillation, network pruning (structured and unstructured), and quantization; (2) Efficient Neural Architecture Design of inherently compact models (e.g., MobileNet, ShuffleNet, EfficientNet series); and (3) Hardware-aware Optimization and Adaptation, encompassing operator fusion, dedicated inference engines, and leveraging heterogeneous systems. ​Results and Conclusion: The analysis highlights key achievements in reducing model size, complexity, and latency while maintaining accuracy. However, fundamental challenges persist, including the accuracy-efficiency tradeoff, hardware fragmentation, the memory wall bottleneck, and privacy/security concerns during deployment. Emerging solutions like neural-symbolic learning, adaptive federated learning, hardware-aware Neural Architecture Search (NAS), Processing-in-Memory (PIM) accelerators, and cross-stack co-design frameworks represent promising future directions. Overcoming these challenges is strategically vital for unlocking the full potential of ubiquitous, real-time edge intelligence.

References

[1] Wang Q, Jin G, Li Q, Wang K, Yang Z, Wang H. Industrial Edge Computing: Vision and Challenges. Information and Control, 2021, 50(3): 257-274.

[2] Hussain H, Tamizharasan P S, Rahul C S. Design possibilities and challenges of DNN models: a review on the perspective of end devices. Artificial Intelligence Review, 2022, 55: 5109–5167.

[3] Hadidi R, Cao J, Xie Y, Asgari B, Krishna T, Kim H. Characterizing the Deployment of Deep Neural Networks on Commercial Edge Devices. IEEE International Symposium on Workload Characterization, 2019: 35–48.

[4] Choudhary T, Mishra V, Goswami A, et al. A comprehensive survey on model compression and acceleration. Artificial Intelligence Review, 2020, 53: 5113–5155.

[5] Reed R. Pruning algorithms—a survey. IEEE Transactions on Neural Networks, 1993, 4(5): 740–747.

[6] Pham-Quoc C, Nguyen X Q, Thinh T N. Towards an FPGA- targeted Hardware/Software Co-design Framework for CNN- based Edge Computing. Mobile Networks and Applications, 2022, 27: 2024–2035.

[7] Li B. Software Framework for Embedded Neural Networks. In: Embedded Artificial Intelligence. Springer, 2024.

[8] Gou J, Yu B, Maybank S J, et al. Knowledge Distillation: A Survey. International Journal of Computer Vision, 2021, 129: 1789–1819.

[9] Yeom SK, Seegerer P, Lapuschkin S, Binder A, Wiedemann S, Müller KR, Samek W. Pruning by explaining: A novel criterion for deep neural network pruning. Pattern Recognition. 2021 Jul 1;115:107899.

[10] He Y, Xiao L. Structured Pruning for Deep Convolutional Dean&Francis ISSN 2959-6157 Neural Networks: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(5): 2900–2919.

[11] Anwar S, Hwang K, Sung W. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems (JETC). 2017 Feb 9;13(3):1-8.

[12] He Y, Kang G, Dong X, Fu Y, Yang Y. Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks. Proceedings of the 27th International Joint Conference on Artificial Intelligence, 2018: 2234–2240.

[13] Zhang X, Zhou X, Lin M, Sun J. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on computer vision and pattern recognition 2018 (pp. 6848-6856).

[14] Liu X, Xu W, Wang Q, Zhang M. Energy-Efficient Computing Acceleration of Unmanned Aerial Vehicles Based on a CPU/FPGA/NPU Heterogeneous System. IEEE Internet of Things Journal, 2024, 11(16): 27126–27138.

[15] Tan T, Cao G. Deep Learning Video Analytics Through Edge Computing and Neural Processing Units on Mobile Devices. IEEE Transactions on Mobile Computing, 2023, 22(3): 1433–1448.

[16] Shuvo M M H, Islam S K, Cheng J, Morshed B I. Efficient Acceleration of Deep Learning Inference on Resource- Constrained Edge Devices: A Review. Proceedings of the IEEE, 2023, 111(1): 42–91.

[17] Yang P, Wang H, Yang J, Qian Z, Zhang Y, Lin X. Deep Learning Approaches for Similarity Computation: A Survey. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(12): 7893–7912.

Downloads

Published

2025-08-26