The Comprehensive Review on Prompt Injection Attacks and Defense Mechanisms in Large Language Models
DOI:
https://doi.org/10.61173/390f5h97Keywords:
Large Language Models, Prompt Injection Attacks, Defense Mechanisms, GCG Algorithm, Semantic Manipulation, Resource Exploitation, Adaptive Defense, CybersecurityAbstract
This review analyzes prompt injection attacks in large language models (LLMs) from 2019 to 2025, addressing critical security challenges as models like ChatGPT proliferate across sectors. We synthesize advances in detection, classification, and mitigation strategies, proposing a tripartite framework categorizing attacks by vector (text/image/speech), mechanism (semantic manipulation, resource exploitation), and impact (data breaches, privacy theft). Key attack vectors include the GCG algorithm, DAN jailbreaks, and resource-exhaustion tactics (e.g., Engorgio). Current defenses are evaluated for efficacy, highlighting scalability gaps and trade-offs between security and model utility. Future priorities include adaptive defense systems leveraging reinforcement learning, interdisciplinary collaboration to address ethical-technical intersections, and open threat intelligence networks for proactive vulnerability management. This work equips researchers and practitioners with actionable strategies to secure LLM ecosystems against evolving adversarial threats.
References
[1] SecureNexusLab LLM-Attack Committee. (2024). Large Language Model Prompt Attack Handbook. SecureNexusLab.
[2] Zou, A., et al. (2023). Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv.
[3] Yong, Z., et al. (2023). Low-Resource Languages Jailbreak GPT-4. CCS.
[4] Dong, Y., et al. (2025). Engorgio: Resource-Exhaustion Attacks on LLM Serving Systems. ICLR.
[5] Fu, J., et al. (2024). Stealthy Data Extraction via Indirect Prompt Injection in Retrieval-Augmented Generation. USENIX Security.
[6] Deng, Y., et al. (2024). Black-Box Prompt Injection via Adversarial Transfer Learning. NDSS.
[7] IBM. (2024, April). What is a prompt injection attack?
[8] OWASP. (2023, October). OWASP Top 10 for LLM Applications (Version 1.1).
[9] Liu, K. (2023, February). The entire prompt of Microsoft Bing Chat? [Blog post].
[10] Wei, A., Haghtalab, N., & Steinhardt, J. (2024). Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems, 36.
[11] Goodside, R. (2022, September). Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions [GitHub].
[12] Sar, E. [omarsar]. (2023, March). Prompt-engineeringguide/guides/prompts-adversarial.md.
[13] Fabrega, A., Namavari, A., Agarwal, R., Nassi, B., & Ristenpart, T. (2024). Exploiting leakage in password managers via injection attacks. 33rd USENIX Security Symposium, 4337– 4354.
[14] Chung, J., Hyun, S., & Heo, J.-P. (2024). Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer. Proceedings of the IEEE/ CVF Conference on Computer Vision and Pattern Recognition, 8795–8805.
[15] Wang, H., Xing, P., Huang, R., Ai, H., Wang, Q., & Bai, X. (2024). InstantStyle-Plus: Style transfer with content-preserving in text-to-image generation. arXiv:2407.00788.
[16] Enono, Paling, Oriettaxx, & Throwawayadvsec. (2023, April). ChatGPT grandma exploit [Forum post].
[17] Wang, Z. A. (2023, September). From DAN to universal prompts: LLM jailbreaking. Deepgram.
[18] Yeung, K., & Ring, L. (2024, March). HiddenLayer research: Prompt injection attacks on LLMs.
[19] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 79–90). ACM.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
