Research and Analysis of Action Recognition Based on Video

Authors

  • Yuexin Dai

DOI:

https://doi.org/10.61173/fvwsry08

Keywords:

Video action recognition, Deep learning, Convolutional neural network, Transformer, Computer vision

Abstract

In recent years, along with the swift development of computer vision and deep learning technologies, videobased action recognition has turned into one of the core research directions within the field of artificial intelligence. Its accomplishments are extensively applied in such real scenarios as intelligent monitoring, human-computer interaction, and autonomous driving. This paper first presents the research background and practical significance of video-based action recognition in recent years then analyzes the current challenges such as those of complex background motion blur and similarity among categories. Next it expounds upon the principle’s structures and application effects of representative models such as 3D convolutional neural networks two - stream networks and models based on the Transformer and also analyzes the advantages and disadvantages of various models. Finally, it summarizes the application scenarios of video action recognition, explores the existing technical difficulties, and also looks forward to the development trends in the future like lightweight models and few - shot learning. This paper offers comprehensive references for researchers in relevant fields, making them able to grasp the research current situation and carry out in - depth research.

References

[1] Feichtenhofer C, Fan H, Malik J, et al. SlowFast Networks for Video Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 43(11): 3586-3597. DOI: 10.1109/TPAMI.2020.2983687.

[2] Bertasius G, Wang H, Torresani L. Video Swin Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(10): 11972-11985. DOI: 10.1109/ TPAMI.2022.3193877.

[3] Sahoo S P. Human Action Recognition Based on Analysis of Video Sequences. Rourkela: National Institute of Technology Rourkela, 2021.

[4] Zhang Y, Li X, Wang L, et al. Light3D: A Lightweight 3D CNN for Edge-Device Video Action Recognition. IEEE Internet of Things Journal, 2023, 10(17): 15218-15228. DOI: 10.1109/ JIOT.2023.3289456. Dean&Francis Yuexin Dai

[5] Girdhar R, Carreira J, Doersch C, et al. ViViT: A Video Vision Transformer.Proceedings of the International Conference on Machine Learning. PMLR, 2021: 3202-3212. DOI: 10.48550/ arXiv.2103.15691.

[6] Misra I, Shrivastava A, Gupta A, et al. TimeSformer: Is Space-Time Attention All You Need for Video Understanding?. International Journal of Computer Vision, 2022, 130(8): 2083- 2101. DOI: 10.1007/s11263-022-01643-4.

[7] Tan M, Le Q V. X3D: Expanding Architectures for Efficient Video Recognition.Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021: 13708-13718. DOI: 10.1109/CVPR46437.2021.01358.

[8] Chen M, Huang X L, Wu M. Application of Lightweight 3D CNN in Action Recognition for Intelligent Monitoring . Control and Decision, 2023, 38 (7): 1765-1772. DOI: 10.13195/ j.kzyjc.2022.0867.

[9] Zhang S Y, Li Y F, Wang Z Y, et al. PoseConv3D: A Skeletal Action Recognition Method Based on 3D Convolutional Neural Network. Journal of Multimedia, 2025, 27 (3): 1865-1878. DOI: 10.1109/TMM.2024.3468921.

[10] Wang L, Zhao X, Liu C. Research on Video Action Recognition Based on Two-Stream Attention Fusion. Pattern Recognition and Artificial Intelligence, 2022, 35 (4): 335-343. DOI: 10.16451/j.cnki.issn1003-6059.202204007.

[11] Li J, Chen Y, Zhang S, et al. Few-Shot Video Action Recognition with Meta-Transformer. Proceedings of the European Conference on Computer Vision. 2024: 456-473. DOI: 10.48550/arXiv.2403.12157.

Downloads

Published

2026-02-28