The Adversarial Attacks towards Human Action Recognition Models: A Comparison between unimodal and multimodal Models

Authors

  • Jizheng Li

DOI:

https://doi.org/10.61173/0r0nva44

Keywords:

FGSM, Adversarial Attacks, Human-Action Recognition, Deep Learning

Abstract

Human-action recognition models are neural networks that analyse visual inputs and provide classification or text outputs. This technology has and will significantly impact society in security, education, healthcare, etc. However, Human-Action Recognition models, like other neural networks, are still susceptible to malicious adversarial attacks. Therefore, this paper proposes an experimental adversarial attack towards ResNet-18 using FGSM. First, ResNet-18 is finetuned using the UCF-101 dataset, and keyframes are selected from sample videos. The keyframes will be given to ResNet-18 for classification while FGSM will be implemented, and ResNet will do another classification of the attached sample. The classification results (Original and Attacked) are given to the language model (GPT-4o) through a prompt that provides the language model with a specific role (e.g. a smart home assistant), and this section is regarded as unimodal. The original and attacked frames will be sent directly instead of the labels in the multimodal section. Lastly, this paper proposes to observe the effects on textual responses generated based on a given prompt and the classification result and evaluate the impact of the attack through Cosine Similarities and Human Evaluation.

References

[1] Haldar, S. (2020, April 9). Gradient-based Adversarial Attacks : An Introduction. Medium. https://medium.com/ swlh/gradient-based-adversarial-attacks-an-introduction- 526238660dc9

[2] Liu, J., Wang, X., & Li, Y. (2021). Understanding the limitations of traditional action recognition methods. Journal of Visual Communication and Image Representation, 79, 103107.

[3] Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.

[4] Moosavian, M. A., & Sadeghi, M. (2020). A comprehensive review of the recent advances in HAR. Computer Vision and Image Understanding, 195, 102924.

[5] Zhang, H., Wu, Y., & Zhang, D. (2021). Adversarial training for deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 32(1), 21-26.

[6] Li, Y., Cheng, S., Wang, Q., & Li, Y. (2022). Keyframe extraction methods for video summarization: a survey. ACM Computing Surveys, 54(5), 1-36.

[7] Gao, S., Jia, X., Ren, X., Tsang, I., & Guo, Q. (2024). Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory. In ArXiv. Sensen Gao, Xiaojun Jia, Xuhong Ren, Ivor Tsang, Qing Guo. https://arxiv.org/pdf/2403.12445

[8] UCF101 - Action Recognition Data Set. (2013, October 17). UCF Centre for Research in Computer Vision. https://www.crcv. ucf.edu/data/UCF101.php

[9] resnet18 — Torchvision main documentation. (n.d.). Retrieved September 3, 2024, from https://pytorch.org/vision/ main/models/generated/torchvision.models.resnet18.html

[10] Loshchilov, I., & Hutter, F. (2019). DECOUPLED WEIGHT DECAY REGULARIZATION. In arXiv. arXiv. https://arxiv.org/pdf/1711.05101

[11] [Pykes, K. (2024, August). Cross-Entropy Loss Function in Machine Learning: Enhancing Model Accuracy. DataCamp. https://www.datacamp.com/tutorial/the-cross-entropy-lossfunction-in-machine-learning

[12] HUMAN ACTIVITY RECOGNITION MODELS USING DEEP RESIDUAL NETWORKS. (n.d.).https://idrlib.iitbhu. ac.in/xmlui/bitstream/handle/123456789/1480/Chapter_5. pdf?isAllowed=y&sequence=14

[13] Adversarial Example Generation — PyTorch Tutorials 2.4.0+cu121 documentation. (n.d.). Retrieved September 3, 2024, from https://pytorch.org/tutorials/beginner/fgsm_tutorial. html

[14] ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. (n.d.). In ScienceDirect. Partha Pratim Ray. https://www. sciencedirect.com/science/article/pii/S266734522300024X

[15] OpenAI. (2023). GPT-4 Technical Report. OpenAI. https:// cdn.openai.com/papers/gpt-4.pdf

[16] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2023). Attention is All You Need. In arXiv. arXiv. https://arxiv.org/pdf/1706.03762

Downloads

Published

2024-10-29