A Comparative Analysis of Diffusion-Based Models in Text-to-Image Generation
DOI:
https://doi.org/10.61173/jttaew34Keywords:
Generative artificial intelligence, Text-to-image generation, Diffusion model, Midjourney, DALL-E 2Abstract
The recent years have seen the rapid development of artificial intelligence image generation methods with the introduction of a wide range of different models to produce images based on a specific text description. The paper explores the design, functionality, and restrictions of four well-known diffusion-based text-to-image generative artificial intelligence systems: Midjourney, DALL-E2, Stable Diffusion, and Imagen. The paper defines the theoretical context of the diffusion probabilistic model and the manner in which it has been applied in the four models of artificial intelligence. This paper uses the comparative and evaluation performance of these models based on the available empirical research studies on the same in relation to image fidelity, prompt adherence, creativity, and bias on various parameters. The results of the analysis show that Stable Diffusion works well in the generation of photorealistic pictures of the human face, Midjourney works well in creativity and visual image in artistic settings, and Imagen works well in complex text description. These models have also exhibited several challenges and limitations. The findings depict the difficulties in the production of precise, just, and imaginative imagery in a variety of situations.
References
[1] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In: Advances in Neural Information Processing Systems. Virtual, December 6-12, 2020, 2020.
[2] Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. Proc IEEE/CVF Conf Comput Vis Pattern Recognit. 2022, 10684-10695.
[3] Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, Ghasemipour SKS, Gontijo Lopes R, Karagol Ayan B, Salimans T, Ho J, Fleet DJ, Norouzi M. Photorealistic textto-image diffusion models with deep language understanding. In: Advances in Neural Information Processing Systems. New Orleans, LA; November 28-December 9, 2022, 2022:36479- 36494.
[4] Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M. Hierarchical text-conditional image generation with CLIP latents. Dean&Francis Yiyang Wu Preprint. Posted online April 12, 2022. arXiv:2204.06125.
[5] Borji A. Generated faces in the wild: quantitative comparison of Stable Diffusion, Midjourney and DALL·E 2. Preprint. Posted online October 2, 2022. arXiv:2210.00586.
[6] Ibrahim I, Abu Talib M, Ammar A, Tabet Aoul KA, Abuimara T. Comparative and experimental analysis of leading text-toimage generative artificial intelligence models for regional residential architectural designs. Results Eng. 2026, 29:108835.
[7] Lan X, An J, Guo Y, et al. Imagining the Far East: exploring perceived biases in AI-generated images of East Asian women. Preprint. Posted online April 8, 2025. arXiv:2504.04865.
[8] Xu M Y. Application of AI Drawing Technology in the Field of Digital Illustration and Its Process. Tomorrow’s Fashion, 2025, (04): 164-166.
[9] Liu F L, Yang L Y. AI Drawing Empowers the Development of 3D Design: Taking Stable Diffusion as an Example. Kunming Metallurgy College Journal, 2025, 41(01): 61-69.
[10] Wang L T. Application of AI Drawing in Textile Pattern Design. Shanghai Apparel, 2024, (12): 19-21.
[11] He J L. Generative AI Drawing Technology Empowers Journalism: Opportunities, Risks, and Mitigation. News Editing, 2024, (05): 125-126.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
