FPGA-Based Implementation and Software Verification of a 3×3 Convolution Core for Edge AI Systems

Authors

  • Hanwen Zhang

DOI:

https://doi.org/10.61173/des73n80

Keywords:

FPGA, convolution, edge computing, Vivado simulation, hardware–software co-verification

Abstract

This study addresses the growing need for lightweight and energy-efficient convolution accelerators in edge artificial intelligence (Edge-AI), where computation is required to occur close to sensors under strict hardware constraints. To support this demand, the paper presents the design and verification of a compact 3×3 convolution core implemented on an FPGA. The research focuses on establishing a reproducible hardware–software coverification workflow that does not rely on physical FPGA boards, which is particularly valuable for academic environments and early-stage prototyping. The convolution module was described in Verilog and functionally validated through behavioral simulation in Xilinx Vivado. In parallel, a Python/NumPy model was developed to replicate the same fixed-point arithmetic, including quantization and ReLU activation. Simulation data exported from Vivado served as the input to the Python verification script. The numerical comparison between the two pipelines demonstrated complete output consistency across all tested pixels, confirming the correctness of the arithmetic pipeline, dataflow control, and activation behavior. The results show that software-driven verification is sufficient to achieve bit-accurate equivalence with the hardware design, significantly reducing development time and improving reproducibility. This workflow provides a practical foundation for future research on scalable FPGAbased convolution accelerators for embedded AI systems.

References

[1] P. Toupas, A. Montgomerie-Corcoran, C.-S. Bouganis and Dean&Francis Hanwen Zhang D. Tzovaras, ‚HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices,‘ IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2023, pp. 144-154, doi: 10.1109/FCCM57271.2023.00024.

[2] E. Wang, J. J. Davis and P. Y. K. Cheung, ‚A PYNQ-Based Framework for Rapid CNN Prototyping,‘ IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2018, pp. 223-223, doi: 10.1109/ FCCM.2018.00057.

[3] J. Xu et al., ‚ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training Quantization,‘ IEEE Transactions on Circuits and Systems I: Regular Papers, vol. 70, no. 12, pp. 5380-5393, Dec. 2023, doi: 10.1109/ TCSI.2023.3315299.

[4] L. D. McLaughlin, L. H. Crockett and R. W. Stewart, ‚A New Design Workflow for PYNQ Enabled Xilinx Platforms Utilising the Simulink Environment for Vivado IPI Abstraction,‘ IEEE 30th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), 2022, pp. 1-1, doi: 10.1109/FCCM53951.2022.9786083.

[5] K. Tajiri and T. Maruyama, ‚FPGA Acceleration of a Composite Kernel SVM for Hyperspectral Image Classification,‘ IEEE Access, vol. 11, pp. 214-226, 2023, doi: 10.1109/ ACCESS.2022.3230066.

[6] D. Suárez, V. Fernández, H. Posadas and P. Sánchez, ‚Accelerating the Verification of Forward Error Correction Decoders by PCIe FPGA Cards,‘ IEEE Embedded Systems Letters, vol. 15, no. 3, pp. 157-160, Sept. 2023, doi: 10.1109/ LES.2022.3218289.

Downloads

Published

2026-02-28