Research on Efficient CNN Acceleration Through Mixed Precision Quantization: A Comprehensive Methodology

计算机科学现场可编程门阵列卷积神经网络量化（信号处理）计算机工程硬件加速边缘设备算法计算机硬件并行计算计算机体系结构嵌入式系统人工智能云计算操作系统

作者

Yizhi He,Wenlong Liu,Muhammad Tahir,Zhao Li,Shaoshuang Zhang,Hussain Bux Amur

出处

期刊：International Journal of Advanced Computer Science and Applications [Science and Information Organization]
日期：2023-01-01 卷期号：14 (12)

链接

thesai.org thesai.orgdoi.org

标识

DOI：10.14569/ijacsa.2023.0141282

摘要

To overcome challenges associated with deploying Convolutional Neural Networks (CNNs) on edge computing devices with limited memory and computing resources, we propose a mixed-precision CNN calculation method on a Field Programmable Gate Array (FPGA). This approach involves a collaborative design encompassing both software and hardware aspects. Initially, we devised a CNN quantization method tailored for the fixed-point operation characteristics of FPGA, addressing the computational challenges posed by floating-point parameters. We introduce a bit-width strategy search algorithm that assigns bit-widths to each layer based on CNN loss variation induced by quantization. Through retraining, this strategy mitigates the degradation in CNN inference accuracy. For FPGA acceleration design, we employ a flow processing architecture with multiple Processing Elements (PEs) to support mixed-precision CNNs. Our approach incorporates a folding design method to implement shared PEs between layers, significantly reducing FPGA resource usage. Furthermore, we designed a data reading method, incorporating a register set buffer between memory and processing elements to alleviate issues related to mismatched data reading and computing speeds. Our implementation of the mixed-precision ResNet20 model on the Kintex-7 Eco R2 development board achieves an inference accuracy of 91.68% and a computing speed 4.27 times faster than the Central Processing Unit (CPU) on the CIFAR-10 dataset, with an accuracy drop of only 1.21%. Compared to a unified 16-bit FPGA accelerator design method, our proposed approach demonstrates an 89-fold increase in computing speed while maintaining similar accuracy.

求助该文献

最长约 10秒，即可获得该文献文件

Research on Efficient CNN Acceleration Through Mixed Precision Quantization: A Comprehensive Methodology

今日热心研友