
Zhuang CAO, Tian ZHANG, Xiaowei HE, Sheng LIU, Junzhong SHEN. Efficient mapping and flexible interconnects: accelerating 3D CNN-based lung nodule segmentation and classification on multi-FPGA[J]. Journal of Zhejiang University Science C, 2026, 27(6): 1-16.
@article{title="Efficient mapping and flexible interconnects: accelerating 3D CNN-based lung nodule segmentation and classification on multi-FPGA",
author="Zhuang CAO, Tian ZHANG, Xiaowei HE, Sheng LIU, Junzhong SHEN",
journal="Journal of Zhejiang University Science C",
volume="27",
number="6",
pages="1-16",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2025.0186"
}
%0 Journal Article
%T Efficient mapping and flexible interconnects: accelerating 3D CNN-based lung nodule segmentation and classification on multi-FPGA
%A Zhuang CAO
%A Tian ZHANG
%A Xiaowei HE
%A Sheng LIU
%A Junzhong SHEN
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 6
%P 1-16
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2025.0186
TY - JOUR
T1 - Efficient mapping and flexible interconnects: accelerating 3D CNN-based lung nodule segmentation and classification on multi-FPGA
A1 - Zhuang CAO
A1 - Tian ZHANG
A1 - Xiaowei HE
A1 - Sheng LIU
A1 - Junzhong SHEN
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 6
SP - 1
EP - 16
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2025.0186
Abstract: three-dimensional convolutional neural networks (3D CNNs) show considerable promise for lung nodule detection. However, their high computational complexity and memory demands present substantial challenges for acceleration on a single field-programmable gate array (FPGA). To address this, we propose efficient mapping schemes for a multi-FPGA platform, leveraging its massive parallelism to maximize computational efficiency. Our system, integrating six customized FPGA boards, achieves state-of-the-art performance, delivering approximately 15.9 tera operations per second (TOPS) for nodule segmentation and approximately 3.8 TOPS for nodule classification. Compared to a central processing unit baseline, it achieves a 128.2× speedup while exhibiting 6.7× higher energy efficiency than a graphics processing unit implementation. Furthermore, the system attains a state-of-the-art recall rate of 87.1% on the real-world clinical benchmark.
[1]Abadi M, Barham P, Chen JM, et al., 2016. TensorFlow: a system for large-scale machine learning. Proc 12th USENIX Conf on Operating Systems Design and Implementation, p.265-283.
[2]Aydonat U, O’Connell S, Capalija D, et al., 2017. An OpenCLTM deep learning accelerator on Arria 10. Proc ACM/SIGDA Int Symp on Field-Programmable Gate Arrays, p.55-64.
[3]Basalama S, Sohrabizadeh A, Wang J, et al., 2023. FlexCNN: an end-to-end framework for composing CNN accelerators on FPGA. ACM Trans Reconfig Technol Syst, 16(2):23.
[4]Chen HX, Song MC, Zhao JC, et al., 2019. 3D-based video recognition acceleration by leveraging temporal locality. Proc 46th Int Symp on Computer Architecture, p.79-90.
[5]Cheng X, 2025. FPGA-based accelerator for convolutional neural networks. Proc Int Conf on Digital Analysis and Processing, Intelligent Computation, p.10-14.
[6]Cui N, Wu YJ, Xin GJ, et al., 2025. Application of quantitative interpretability to evaluate CNN-based models for medical image classification. IEEE Access, 13:89386-89398.
[7]Dey R, Lu ZJ, Hong Y, 2018. Diagnostic classification of lung nodules using 3D neural networks. Proc IEEE 15th Int Symp on Biomedical Imaging, p.774-778.
[8]Diaconu D, Lin X, Blott M, et al., 2026. A survey of FPGA-based 3D CNN accelerators and hardware-aware algorithmic optimizations. NACM Comput Surv, 58(6):155.
[9]Fan HX, Luo C, Zeng CL, et al., 2019. F-E3D: FPGA-based acceleration of an efficient 3D convolutional neural network for human action recognition. Proc IEEE 30th Int Conf on Application-Specific Systems, Architectures and Processors, p.1-8.
[10]Fan HX, Liu SL, Que ZQ, et al., 2023. High-performance acceleration of 2-D and 3-D CNNs on FPGAs using static block floating point. IEEE Trans Neur Netw Learn Syst, 34(8):4473-4487.
[11]Fowers J, Ovtcharov K, Papamichael M, et al., 2018. A configurable cloud-scale DNN processor for real-time AI. Proc ACM/IEEE 45th Annual Int Symp on Computer Architecture, p.1-14.
[12]Gao L, Luo ZQ, Wang L, 2025. Convolutional neural network acceleration techniques based on FPGA platforms: principles, methods, and challenges. Information, 16(10):914.
[13]Geng T, Wang TQ, Sanaullah A, et al., 2018. FPDeep: acceleration and load balancing of CNN training on FPGA clusters. Proc IEEE 26th Annual Int Symp on Field-Programmable Custom Computing Machines, p.81-84.
[14]Hegde K, Agrawal R, Yao YL, et al., 2018. Morph: flexible acceleration for 3D CNN-based video understanding. Proc 51st Annual IEEE/ACM Int Symp on Microarchitecture, p.933-946.
[15]Huang G, Liu Z, Van Der Maaten L, et al., 2017. Densely connected convolutional networks. Proc IEEE Conf on Computer Vision and Pattern Recognition, p.2261-2269.
[16]Huang XJ, Shan JJ, Vaidya V, 2017. Lung nodule detection in CT using 3D convolutional neural networks. Proc IEEE 14th Int Symp on Biomedical Imaging, p.379-383.
[17]Isensee F, Jaeger PF, Kohl SAA, et al., 2021. nnU-Net: a self-configurationuring method for deep learning-based biomedical image segmentation. Nat Methods, 18(2):203-211.
[18]Jiang JY, Zhou YA, Gong YH, et al., 2025. FPGA-based acceleration for convolutional neural networks: a comprehensive review.
[19]Jiang WW, Sha EHM, Zhang XY, et al., 2019. Achieving super-linear speedup across multi-FPGA for real-time DNN inference. ACM Trans Embed Comput Syst, 18(5s):67.
[20]Khan FH, Pasha MA, Masud S, 2023. Towards designing a hardware accelerator for 3D convolutional neural networks. Comput Electr Eng, 105:108489.
[21]Kuang HL, Wang YH, Tan XZ, et al., 2025. LW-CTrans: a lightweight hybrid network of CNN and Transformer for 3D medical image segmentation. Med Image Anal, 102:103545.
[22]Lin CY, Guo SM, Lien JJJ, 2024. Development of a modified 3D region proposal network for lung nodule detection in computed tomography scans: a secondary analysis of lung nodule datasets. Cancer Imaging, 24(1):40.
[23]Litjens G, Kooi T, Bejnordi BE, et al., 2017. A survey on deep learning in medical image analysis. Med Image Anal, 42:60-88.
[24]Liu SL, Fan HX, Ferianc M, et al., 2022. Toward full-stack acceleration of deep convolutional neural networks on FPGAs. IEEE Trans Neur Netw Learn Syst, 33(8):3974-3987.
[25]Lu YH, Qi XY, Li Y, et al., 2024. Automatic implementation of large-scale CNNs on FPGA cluster based on HLS4ML. Proc IEEE Int Symp on Parallel and Distributed Processing with Applications, p.1080-1087.
[26]Messay T, Hardie RC, Rogers SK, 2010. A new computationally efficient CAD system for pulmonary nodule detection in CT imagery. Med Image Anal, 14(3):390-406.
[27]Microsoft, 2018. Project Catapult. https://www.microsoft.com/en-us/research/project/project-catapult [Accessed on May 10, 2026].
[28]Microsoft, 2019. Project Brainwave. https://www.microsoft.com/en-us/research/project/project-brainwave/ [Accessed on May 10, 2026].
[29]Motamedi M, Gysel P, Akella V, et al., 2016. Design space exploration of FPGA-based deep convolutional neural networks. Proc 21st Asia and South Pacific Design Automation Conf, p.575-580.
[30]NVIDIA, 2019. NVIDIA cuDNN. https://developer.nvidia.com/cudnn [Accessed on May 10, 2026].
[31]Qin JJ, Xiong J, Liang ZT, 2025. CNN-Transformer gated fusion network for medical image super-resolution. Sci Rep, 15(1):15338.
[32]Ronneberger O, Fischer P, Brox T, 2015. U-Net: convolutional networks for biomedical image segmentation. Proc 18th Int Conf on Medical Image Computing and Computer-Assisted Intervention, p.234-241.
[33]Shen JZ, Qiao Y, Huang Y, et al., 2018a. Towards a multi-array architecture for accelerating large-scale matrix multiplication on FPGAs. Proc IEEE Int Symp on Circuits and Systems, p.1-5.
[34]Shen JZ, Huang Y, Wang ZL, et al., 2018b. Towards a uniform template-based architecture for accelerating 2D and 3D CNNs on FPGA. Proc ACM/SIGDA Int Symp on Field-Programmable Gate Arrays, p.97-106.
[35]Shen JZ, Wang DG, Huang Y, et al., 2019. Scale-out acceleration for 3D CNN-based lung nodule segmentation on a multi-FPGA system. Proc 56th Annual Design Automation Conf, Article 207.
[36]Shen JZ, Huang Y, Wen M, et al., 2020. Toward an efficient deep pipelined template-based architecture for accelerating the entire 2-D and 3-D CNNs on FPGA. IEEE Trans Comput-Aided Des Integr Circ Syst, 39(7):1442-1455.
[37]Siegel RL, Giaquinto AN, Jemal A, 2024. Cancer statistics, 2024. CA Cancer J Clin, 74(1):12-49.
[38]Siegel RL, Kratzer TB, Giaquinto AN, et al., 2025. Cancer statistics, 2025. CA Cancer J Clin, 75(1):10-45.
[39]Tan MXN, Deklerck R, Jansen B, et al., 2011. A novel computer-aided lung nodule detection system for CT images. Med Phys, 38(10):5630-5645.
[40]Varughese DA, Sridevi S, 2026. Optimization strategies and algorithms for accelerating CNN on FPGA: a comprehensive review. Arch Comput Methods Eng, 33(1):1205-1226.
[41]Wang JX, Zhang XW, Tang W, et al., 2025. A multi-view CNN model to predict resolving of new lung nodules on follow-up low-dose chest CT. Insights Imaging, 16(1):138.
[42]Wang YC, Wang Y, Li HW, et al., 2019. Systolic cube: a spatial 3D CNN accelerator architecture for low power video analysis. Proc 56th Annual Design Automation Conf, Article 210.
[43]Wikipedia, 2019. Ipv4. https://en.wikipedia.org/wiki/IPv4 [Accessed on May 10, 2026].
[44]Wu LH, Zhang M, Piao Y, et al., 2025. CNN-Transformer rectified collaborative learning for medical image segmentation. IEEE Trans Circ Syst Video Technol, 35(5):4072-4086.
[45]Zhang C, Wu D, Sun JY, et al., 2016. Energy-efficient CNN implementation on a deeply pipelined FPGA cluster. Proc Int Symp on Low Power Electronics and Design, p.326-331.
[46]Zhang C, Sun GY, Fang ZM, et al., 2019. Caffeine: toward uniformed representation and acceleration for deep convolutional neural networks. IEEE Trans Comput-Aided Des Integr Circ Syst, 38(11):2072-2085.
[47]Zhang WT, Zhang JX, Shen MH, et al., 2019. An efficient mapping approach to large-scale DNNs on multi-FPGA architectures. Proc Design, Automation & Test in Europe Conf & Exhibition, p.1241-1244.
CLC number: TP391.4
On-line Access: 2026-07-29
Received: 2025-12-24
Revision Accepted: 2026-05-15
Crosschecked: 2026-07-29
Cited: 0
Clicked: 4
Open peer comments: Debate/Discuss/Question/Opinion
<1>