A review of 3D reconstruction methods based on deep learning

Authors

  • Liwei Wang

DOI:

https://doi.org/10.61173/vrwd4a81

Keywords:

deep learning, 3D reconstruction, NeRF, 3DGS

Abstract

3D reconstruction is a technical process that constructs a digital 3D model of a target object from low-dimensional data. It plays an important role in medical imaging, cultural relics protection and other fields.Traditional 3D reconstruction techniques suffer from challenges such as difficult feature extraction and heavy manual intervention. Therefore, deep learning has been introduced into this field. After extensive literature review, this paper systematically summarizes classic 3D reconstruction algorithms using deep learning methods, categorizing them into explicit and implicit representation approaches. As cutting-edge technologies in 3D reconstruction, Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) hold significant promise. This paper briefly introduces the fundamental principles and recent advancements of dynamic scenes, outlines commonly used dynamic scene datasets and performance metrics, and compares their performance on the D-NeRF datasets. It concludes by summarizing the main challenges in 3D reconstruction and looks ahead to future developments in technology integration and reducing memory costs for large-scale scenes.

References

[1] Hartley R. Multiple view geometry in computer vision[M]. Cambridge university press, 2003.

[2] Mildenhall B, Srinivasan P P, Tancik M, et al. Nerf: Representing scenes as neural radiance fields for view synthesis[J]. Communications of the ACM, 2021, 65(1): 99-106.

[3] Kerbl B, Kopanas G, Leimkühler T, et al. 3d gaussian splatting for real-time radiance field rendering[J]. ACM Trans. Graph., 2023, 42(4): 139:1-139:14.

[4] Wu Z, Song S, Khosla A, et al. 3d shapenets: A deep representation for volumetric shapes[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2015: 1912-1920.

[5] Choy C B, Xu D, Gwak J Y, et al. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction[C]// Computer vision–ECCV 2016: 14th European conference, amsterdam, the netherlands, October 11-14, 2016, proceedings, part VIII 14. Springer International Publishing, 2016: 628-644.

[6] Xie H, Yao H, Sun X, et al. Pix2vox: Context-aware 3d reconstruction from single and multi-view images[C]// Proceedings of the IEEE/CVF international conference on computer vision. 2019: 2690-2698.

[7] Qi C R, Su H, Mo K, et al. Pointnet: Deep learning on point sets for 3d classification and segmentation[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 652-660.

[8] Yuan W, Khot T, Held D, et al. Pcn: Point completion network[C]//2018 international conference on 3D vision (3DV). IEEE, 2018: 728-737.

[9] Wang N, Zhang Y, Li Z, et al. Pixel2mesh: Generating 3d mesh models from single rgb images[C]//Proceedings of the European conference on computer vision (ECCV). 2018: 52-67.

[10] Wen C, Zhang Y, Li Z, et al. Pixel2mesh++: Multi-view 3d mesh generation via deformation[C]//Proceedings of the IEEE/ CVF international conference on computer vision. 2019: 1042- 1051.

[11] Park J J, Florence P, Straub J, et al. Deepsdf: Learning continuous signed distance functions for shape representation[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 165-174.

[12] Pumarola A, Corona E, Pons-Moll G, et al. D-nerf: Neural radiance fields for dynamic scenes[C]//Proceedings of the IEEE/ CVF conference on computer vision and pattern recognition. 2021: 10318-10327.

[13] Park K, Sinha U, Hedman P, et al. HyperNeRF: a higherdimensional representation for topologically varying neural radiance fields[J]. ACM Transactions on Graphics, 2021, 40(6): Article No.238.

[14] Fang J, Yi T, Wang X, et al. Fast dynamic radiance fields with time-aware neural voxels[C]//SIGGRAPH Asia 2022 Conference Papers. 2022: 1-9.

[15] Fridovich-Keil S, Meanti G, Warburg F R, et al. K-planes: Explicit radiance fields in space, time, and appearance[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023: 12479-12488.

[16] Cao A, Johnson J. Hexplane: A fast representation for dynamic scenes[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023: 130-141.

[17] Yang Z, Gao X, Zhou W, et al. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction[C]// Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 20331-20341.

[18] Wu G, Yi T, Fang J, et al. 4d gaussian splatting for realtime dynamic scene rendering[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 20310-20320.

[19] Lin Y, Dai Z, Zhu S, et al. Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle[C]//Proceedings of the IEEE/ CVF Conference on Computer Vision and Pattern Recognition. 2024: 21136-21145.

[20] Huang Y H, Sun Y T, Yang Z, et al. Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 4220-4230.

Downloads

Published

2025-08-26