A Highly-Parallel AI Accelerator Architecture for Convolution and Activation, Implemented in Verilog
DOI:
https://doi.org/10.61173/fmwnqv14Keywords:
Convolutional neural networks, FPGA acceleration, hardware implementation, LeNet-5, edge computingAbstract
LeNet-5 is a classic Convolutional Neural Network (CNN) model whose core structure, such as the C1 Convolution layer, remains constrained by hardware resources like computing power, power consumption, and storage bandwidth for real-time inference in embedded and edge computing scenarios. To break through this bottleneck and enhance the computing efficiency of artificial intelligence in resource-constrained environments, this study focuses on the design of a dedicated hardware accelerator for the Convolution Layer 1 (C1) of LeNet-5 and its subsequent Rectified Linear Unit (ReLU). A highly parallel convolution computing architecture was constructed, enabling synchronous operation and data reuse across multiple groups of convolution units (CUs), which significantly improved computational throughput and energy efficiency. The experimental results show that while controlling the resource consumption of the Field Programmable Gate Array (FPGA), this accelerator has a significant improvement in inference speed compared with the pure software implementation, successfully verifying the technical feasibility and engineering advantages of hardwareization of the basic operators of Convolution neural networks. This research not only provides a reusable hardware prototype and optimization path for the efficient deployment of CNN models in embedded terminals, but also lays an important theoretical and technical foundation for the industrial application of artificial intelligence edge computing, possessing high academic innovation and practical application value.
References
[1] Liu Z, Dou Y, Jiang J, et al. Throughput-Optimized FPGA [8] Pang M, Wei X, Zhang Y, et al. Efficient Adaptive Accelerator for Deep Convolutional Neural Networks. ACM Convolutional Neural Network Accelerator for Low-Resource Transactions on Reconfigurable Technology and Systems Chips. Computer Science, 2025, 52(04): 94-100. (TRETS), 2017, 10(3): 1-23. [9] Basalama S, Sohrabizadeh A, Wang J, et al. FlexCNN: An
[2] Li H, Cheng C, Juntong Y, et al. Multi-Scale Feature end-to-end framework for composing CNN accelerators on Fusion Convolutional Neural Network for Indoor Small Target FPGA. ACM Transactions on Reconfigurable Technology and Detection. Frontiers in Neurorobotics, 2022, 16881021-881021. Systems, 2023, 16(2): 1-32.
[3] Zhao Z, Yang S, Ma Z. Research on Vehicle Licence Plate [10] Samayoa W F, Crespo M L, Cicuttin A, et al. A survey on Character Recognition Based on the Convolutional Neural FPGA-based heterogeneous clusters architectures. IEEE Access, Network LeNet-5. Journal of System Simulation, 2010, 22(03): 2023, 11: 67679-67706. 638-641. [11] Cong J, Lau J, Liu G, et al. FPGA HLS today: successes,
[4] Yu N, Jiao P, Zheng Y. Handwritten digits recognition based challenges, and opportunities. ACM Transactions on on improved LeNet5//The 27th Chinese control and decision Reconfigurable Technology and Systems (TRETS), 2022, 15(4): conference (2015 CCDC). IEEE, 2015: 4871-4875. 1-42.
[5] Wang Y, Xie K, Chen S, Hu J, Chang S. A universal design [12] Singh G, Alser M, Cali D S, et al. FPGA-based nearon hardware acceleration of convolutional neural networks. memory acceleration of modern data-intensive applications. Computer Engineering & Science, 2023, 45(4): 577-581. IEEE Micro, 2021, 41(4): 39-48.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
