The foundation, current situation and future prospects of pre-training large language models

Authors

  • Haoran Han
  • Siyao Wu
  • Jinyao Yang
  • Yizhuo Zhao

DOI:

https://doi.org/10.61173/yha53v12

Keywords:

Large language models, GPT model, advanced GPT models

Abstract

The field of artificial intelligence has developed rapidly recently, and large language model technology, as a representative technology of it, can provide general knowledge and make many downstream tasks easier and more convenient. However, although many people use large language models to do some work, they still lack a systematically summarized literature. Therefore, in this article, we made a systematic summary. We first wrote about the early large language models, then we presented the development of GPT and how to use the GPT model, then we introduced the advanced GPT models, and finally we mentioned the risks and challenges faced by the GPT model. Our work can help users better use large language models.

References

[1] Kammersgaard J. Four different perspectives on human– computer interaction. International Journal of Man-Machine Studies, 1988, 28(4): 343-362.

[2] Hutchins J. The history of machine translation in a nutshell. Retrieved December, 2005, 20(2009): 1-1.

[3] Feng Z. Formal Models of Neural Network and Deep Learning//Formal Analysis for Natural Language Processing: A Handbook. Singapore: Springer Nature Singapore, 2023: 653- 752.

[4] Radford A, Narasimhan K, Salimans T, et al. Improving language understanding by generative pre-training. 2018.

[5] Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners. OpenAI blog, 2019, 1(8): 9.

[6] Kaplan J, McCandlish S, Henighan T, et al. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.

[7] Brown T, Mann B, Ryder N, et al. Language models are few-shot learners. Advances in neural information processing systems, 2020, 33: 1877-1901.

[8] Achiam J, Adler S, Agarwal S, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.

[9] Rahaman M S, Ahsan M M T, Anjum N, et al. From Dean&Francis ChatGPT-3 to GPT-4: a significant advancement in ai-driven NLP tools. Journal of Engineering and Emerging Technologies, 2023, 2(1): 1-11.

[10] Giray L. Prompt engineering with ChatGPT: a guide for academic writers. Annals of biomedical engineering Prompt engineering for ChatGPT: a quick guide to techniques, tips, 2023, 51(12): 2629-2633.

[11] Ekin S. Prompt engineering for ChatGPT: a quick guide to techniques, tips, and best practices[J]. Authorea Preprints, 2023.

[12] Dai Z, Yang Z, Yang Y, et al. Transformer-xl: Attentive language models beyond a fixed-length context[J]. arXiv preprint arXiv:1901.02860, 2019.

[13] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in neural information processing systems, 2017.

[14] Sutton R S, Barto A G. Reinforcement learning: An introduction. MIT press, 2018.

[15] Cristani M, Tomazzoli C. A multimodal approach to relevance and pertinence of documents//Trends in Applied Knowledge-Based Systems and Data Science: 29th International Conference on Industrial Engineering and Other Applications of Applied Intelligent Systems, IEA/AIE 2016, Morioka, Japan, August 2-4, 2016, Proceedings 29. Springer International Publishing, 2016: 157-168.

[16] Hirasawa T, Kaneko M, Imankulova A, et al. Pre-trained word embedding and language model improve multimodal machine translation: A case study in Multi30K. IEEE Access, 2022, 10: 67653-67668.

[17] Oppenlaender J, Hämäläinen J. Mapping the challenges of HCI: An application and evaluation of ChatGPT and GPT-4 for cost-efficient question answering. arXiv preprint arXiv:2306.05036, 2023.

[18] Chen X, Ye J, Zu C, et al. How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks. arXiv preprint arXiv:2303.00293, 2023.

Downloads

Published

2024-06-06