Image Segmentation Based on Transformer-Class Methods
DOI:
https://doi.org/10.61173/ygcq9v41Keywords:
Transformer, U-Net, Image SegmentationAbstract
Convolutional neural networks (CNNs) have many problems, including the inability to encode long-range details of images, failure to capture the global information, and other problems, like vanishing or exploding gradients during image segmentation tasks. Since its introduction in 1985, the Transformer model has proven to be very beneficial in natural language processing and computer vision. This paper reviews the application and development of Transformer models in image segmentation tasks within the biomedical, industrial manufacturing and agricultural environments in order to promote joint development of Transformer architecture and image segmentation research. The article explains the base and relevance of this study and dwells on the discussion of the three key challenges as follows: increasing contour accuracy, generalization ability and strength, and lowering the computation cost. It is also important to mention that in the future, the Transformer model will be more versatile by optimizing its architecture and algorithms, and decreasing the rate at which the parameters are updated during the process of adapting to new tasks. Moreover, combining the Transformer with additional network architectures is likely to create a dynamic tradeoff between local and long-range dependencies.
References
[1] Long J, Shelhamer E, Darrell T. Fully convolutional [11] Chen H, Xiao Z. Swin-TUNA: A novel PEFT approach for networks for semantic segmentation. Proceedings of the IEEE accurate food image segmentation. 2025. Conference on Computer Vision and Pattern Recognition, 2015: [12] Xu H, Song J, Zhu Y. Evaluation and comparison of 3431–3440. semantic segmentation networks for rice identification based on
[2] Chen L C, Papandreou G, Kokkinos I, et al. Semantic image Sentinel-2 imagery. Remote Sensing, 2023, 15(6): 1499. segmentation with deep convolutional nets and fully connected [13] Wang H, Chen X, Zhang T, et al. CCTNet: Coupled CNN CRFs. 2014. and transformer network for crop segmentation of remote
[3] Chen L C, Papandreou G, Kokkinos I, et al. Deeplab: sensing images. Remote Sensing, 2022, 14(9): 1956. Semantic image segmentation with deep convolutional [14] Xiang J, Liu J, Chen D, et al. CTFuseNet: A multinets, atrous convolution, and fully connected CRFs. IEEE scale CNN–transformer feature fused network for crop type Transactions on Pattern Analysis and Machine Intelligence, segmentation on UAV remote sensing imagery. Remote Sensing, 2017, 40(4): 834–848. 2023, 15(4): 1151.
[4] Chen L C, Papandreou G, Schroff F, et al. Rethinking atrous
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
