Research on the Evolution and Classification of Autoregressive Text-to-Image Generation Models

Authors

  • Xuan Pan
  • Haoxuan Huang

DOI:

https://doi.org/10.61173/mn03wc25

Keywords:

Autoregressive Model, Text-Guided Image Generation, Cross-Modal Alignment, Dynamic Token Mechanism

Abstract

This paper systematically reviews the evolution and classification of auto regressive (AR) models in text-guided image generation, with a focus on four representative technical approaches: ARINAR, Token-Shuffle, SimpleAR, and LlamaGen, summarizing their strengths and limitations. The study shows that AR models, leveraging their element-by-element generation mechanism, excel in controllability and cross-modal alignment, yet remain constrained by challenges such as the trade-off between generation efficiency and resolution, insufficient complex semantic mapping, and high sensitivity during training. To address these issues, three optimization directions are proposed: introducing a dynamic token mechanism and parallel acceleration to improve generation efficiency; enhancing structured semantic modeling to refine complex semantic generation; and optimizing robust training strategies to strengthen model generalization capabilities. This research provides a theoretical foundation and technical reference for the improvement and application of AR models, contributing significantly to advancing the practical development of text-to-image generation technologies and enhancing the real-world implementation of AR-based generative models.

References

[1] Chen, Y., Ma, Z., Jia, G., Jiang, C., Li, J., Zhou, B.: Contextaware autoregressive models for multi-conditional image generation. arXiv preprint, arXiv:2505.12274,1-12(2025).

[2] Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., Liu, X.: Infinity∞: scaling bitwise autoregressive modeling for high-resolution image synthesis. arXiv preprint arXiv:2412.04431,1-14 (2024).

[3] Chen, M., Sutskever, I.: Zero-shot text-to-image generation. In: Proceedings of the 38th International Conference on Machine Learning, pp. 8821–8831. PMLR (2021)

[4] Zou, X., Shao, Z., Yang, H., Wang, J., Peng, Y., Shan, Y., Yang, X., & Zhang, L.: CogView: mastering text - to - image generation via transformers. Advances in Neural Information Processing Systems, 34, 5678 - 5690 (2021).

[5] Wu, C., et al.: Nüwa: visual synthesis pre-training for neural visual world creation. In: Proceedings of the European Conference on Computer Vision (pp. 1–20). Springer, Cham.

[6] Tao, C., et al.: Autoregressive models in vision: A survey. arXiv preprintarXiv:2411.05902, 1-54 (2024).

[7] Zhao, Q., Gould, S., Zheng, L.: ARINAR: bi-level autoregressive feature-by-feature generative models. arXiv preprint arXiv:2503.02883, 1-6 (2025).

[8] Ma, X., Sun, P., Ma, H., et al.: Token-shuffle: towards highresolution image generation with autoregressive models. arXiv preprint arXiv:2504.12281,1-24 (2025).

[9] Wang, J., Tian, Z., Wang, X., et al.: SimpleAR: pushing the frontier of autoregressive visual generation through pretraining, SFT, and RL. arXiv preprint arXiv:2504.11455, 1-12 (2025).

[10] Sun, P., Jiang, Y., Chen, S., et al.: Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 1-26 (2024).

Downloads

Published

2025-12-19