A Survey of Textual Adversarial Attacks and Defenses on Large Language Models (LLMs)
DOI:
https://doi.org/10.61173/m8d1p054Keywords:
Large Language Models, Adversarial Attacks, Prompt Injection, Jailbreaking Attacks, Model SecurityAbstract
The broad use of large language models (LLMs) like GPT-4 and LLaMA in dialogue systems, content generation, etc., causes security flaws of these models to rise. The current study examines progress made in recent years on textual adversarial attacks and defenses for LLMs, focusing on special attack vectors and defensive techniques targeting LLMs. Firstly, a 'Target—Technology—Scenario' three-dimensional attack classification framework that mainly consists of typical kinds of attacks including PIA and JBA. Secondly, defense mechanisms from two different perspectives: alignment enhancement in the training phase and security controls in the inference phase. In addition, by conducting experiments using benchmark datasets (HELM), existing technical drawbacks and future work directions are also discussed. The aim is to offer guidance and references for subsequent work on theoretical studies of LLM security and promoting the building of stronger NLP systems.
References
[1] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
[2] OpenAI. (2023). GPT-4 Technical Report. arXiv preprint arXiv:2303.08774.
[3] Shayegani, M., Mamun, M. S., & Jajodia, S. (2023). Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. ACM Computing Surveys, 56(3), 1–38.
[4] Liu, Y. P., Jia, Y. Q., Geng, R. P., Jia, J. Y., & Gong, N. Z. Q. (2024). Formalizing and Benchmarking Prompt Injection Attacks and Defenses. In Proceedings of the 33rd USENIX Security Symposium (pp. 123–140), Philadelphia, PA, USA.
[5] Liu, Y., Deng, G., Li, Y. K., Zhang, Y., & Wang, X. (2023). Prompt Injection Attack Against LLM-Integrated Applications. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 45–56), Los Angeles, CA, USA.
[6] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 79–90), Los Angeles, CA, USA.
[7] Jha, P., Sharma, A., & Jajodia, S. (2024). LLM Stinger: Jailbreaking LLMs Using RL Fine-Tuned LLMs. IEEE Transactions on Dependable and Secure Computing, 21(5), 4559–4573.
[8] Carlini, N., Jagielski, M., & Tramèr, F. (2024). Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks. In Proceedings of the 41st International Conference on Machine Learning (pp. 3245–3258), Vienna, Austria.
[9] Chen, S. Y., Pfohl, K., Cole-Lewis, H., & Lewis, C. (2024). A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models. Journal of Biomedical Informatics, 154(C), 104644.
[10] Tramèr, F., Boneh, D., & Poovendran, R. (2018). Adversarial Machine Learning. In Handbook of Cyber Security (pp. 1–28). Springer.
[11] Zhang, H., Yuan, X., & Li, X. (2020). A Survey of Adversarial Attacks and Defenses in Deep Learning. IEEE Access, 8, 151613–151633.
[12] Chen, X., Liu, X., & Li, Y. (2023). Backdoor Attacks for In-Context Learning with Language Models. In Proceedings of the 40th International Conference on Machine Learning (pp. 2563–2575), Honolulu, HI, USA.
[13] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., ... & Amodei, D. (2022). Training Language Models to Follow Instructions with Human Feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.
[14] Anthropic. (2024). Many-Shot Jailbreaking: Exploiting Long Context Windows in LLMs. Anthropic Research Blog. https://www.anthropic.com/research/many-shotjailbreaking
[15] Hao, Y., & Liu, X. G. (2024). InjeCGuard: Benchmarking and Mitigating Over-Defense in Prompt Injection Guardrail Models. IEEE Transactions on Artificial Intelligence, 5(2), 1–14.
[16] Wang, Z., Li, Y., & Zhang, Y. (2024). Stop Reasoning! When Multimodal LLMs with Chain-of-Thought Reasoning Meets Adversarial Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1234–1243), Seattle, WA, USA.
[17] Wang, L., Li, Y., & Zhang, Y. (2023). Holistic Evaluation of Language Models. Machine Learning and Systems, 5, 1–14.
[18] Li, J., Liu, X., & Li, Y. (2024). Interpretable Attacks on Large Language Models via Attention Visualization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 1234–1243), Singapore.
[19] Liu, X., Zhang, Y., & Li, Y. (2024). Federated Learning for LLM Security: A Case Study in Healthcare. IEEE Journal of Biomedical and Health Informatics, 28(5), 1–14.
[20] Google AI. (2023). Secure AI Framework (SAIF): A New Approach to AI Safety. IEEE Security & Privacy, 21(6), 12–21.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
