Low-Power Hardware Implementation and Evaluation of Sparse CNN for MNIST
DOI:
https://doi.org/10.61173/azhd3v80Keywords:
Sparse CNN, Zero-skip, Clock enable, VCD-based power estimation, Low-power acceleratorAbstract
To address the low-power requirements of edge computing, this work takes MNIST classification as a case study and investigates a sparsity-driven lightweight CNN hardware implementation. Under a unified toolchain, we evaluate the gate-level power of the first convolution–pooling block (Conv1) using VCD-based switching-activity analysis. A sparsity sweep is conducted to study the energy–accuracy trade-offs of zero-skipping. On a 500-image evaluation subset, the baseline achieves 98.0% accuracy with 0.236 μW power. Increasing sparsity to 65% reduces Conv1 power to 0.164 μW (30.5% lower) with 96.0% accuracy. When accuracy must remain at the 98.0% baseline level, a 45% sparsity setting yields 0.182 μW power. The comparison between Zero-Skip and Zero-Skip+CE further indicates that the current CE implementation behaves as data gating rather than true clock gating, providing no additional energy benefit.. We provide detailed descriptions of model configuration, sparsity and gating implementation, VCD statistics and power-conversion settings required for reproducible experiments, thereby offering quantitative evidence and engineering guidelines for sparsity selection and energy-efficiency optimization of lightweight CNNs on resource-constrained platforms.
References
[1] Han S, Mao H, Dally W J. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In: Proc. Int. Conf. Learning Representations (ICLR), 2016.
[2] Blalock D, Gonzalez Ortiz J J, Frankle J, Guttag J. What is the state of neural network pruning?. Proc. Machine Learning and Systems (MLSys), 2020, 2, 129-146.
[3] Véstias M P, Duarte R P, de Sousa J T, Neto H C. Fast convolutional neural networks in low density FPGAs using zeroskipping and weight pruning. Electronics, 2019, 8(11), 1321.
[4] Han S, Liu X, Mao H, Pu J, Pedram A, Horowitz M A, Dally W J. EIE: Efficient inference engine on compressed deep neural network. In: Proc. 43rd Int. Symp. Computer Architecture (ISCA), 2016, 243-254.
[5] Parashar A, Rhu M, Mukkara A, Puglielli A, Venkatesan R, Khailany B, Emer J S, Keckler S W, Dally W J. SCNN: An accelerator for compressed-sparse convolutional neural networks. In: Proc. 44th Annu. Int. Symp. Computer Architecture (ISCA), 2017, 27-40.
[6] Aimar A, Mostafa H, Calabrese E, Rios-Navarro A, Tapiador- Morales R, Lungu I A, Milde M B, Corradi F, Linares-Barranco A, Liu S C, Delbruck T. NullHop: A flexible convolutional neural network accelerator based on sparse representations of feature maps. IEEE Trans. Neural Netw. Learn. Syst., 2019, 30(3), 644-656.
[7] Chen Y H, Krishna T, Emer J S, Sze V. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks. In: Proc. 43rd Int. Symp. Computer Architecture (ISCA), 2016, 367-379.
[8] Jacob B, Kligys S, Chen B, Zhu M, Tang M, Howard A, Adam H, Kalenichenko D. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2018, 2704-2713.
[9] LeCun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proc. IEEE, 1998, 86(11), 2278-2324.
[10] Rabaey J M, Chandrakasan A, Nikolić B. Digital Integrated Circuits: A Design Perspective. 2nd ed. Prentice Hall, 2003.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
