Endangered Species Recognition from Camera Trap Images with Vision Transformers

Authors

  • Yunke Wang

DOI:

https://doi.org/10.61173/9dt9q857

Keywords:

Endangered species, Vision Transformer, CNN, Deep learning, Grad-CAM, Test-time augmentation.

Abstract

Automatic identification of endangered species from camera trap images grows more important for wildlife conservation. This task still poses challenges. Lighting conditions vary. Partial occlusion occurs. There are subtle inter-species differences. These differences require fine visual discrimination. This study proposes a dualbranch deep learning ensemble model. It integrates Swin Transformer and ConvNeXt architectures. The model captures complementary global and local features. It’s for the identification of 10 rare species. The team used 2,000 research-grade images. These images come from iNaturalist. Our model reached a Top-1 accuracy of 90.83%. That’s 3.5% higher than EfficientNet-B0, ViT-B16, and Swin-T. The training process used progressive unfreezing. It also used layer-wise learning rate decay. These methods achieve stable multi-scale feature adaptation on limited data. They also suppress overfitting. GradCAM visualizations confirm that the model consistently attends to anatomically discriminative regions—such as rosette patterns and stripe configurations—thereby reducing interspecies confusion. Test-time augmentation further enhances robustness against occlusion and illumination variability. The final system supports practical edge deployment, running at 15 FPS on an NVIDIA Jetson Nano with INT8 quantization. This work demonstrates that hybrid Transformer–CNN architectures are effective and deployable for real-world conservation monitoring.

References

[1] Norouzzadeh M S, Nguyen A, Kosmala M, Swanson A, Clune M S, Clune J. A large-scale benchmark for wildlife identification from camera trap images. IEEE Winter Conference on Applications of Computer Vision, 2021: 662-673.

[2] Tabak M A, Miller M J, Thompson A K, Hinton H J, Share K C, McClaim L P. Machine learning to classify animal species in camera trap images: Applications in ecology. Methods in Ecology and Evolution, 2019, 10(4): 585-590.

[3] Gomez-Villa C, Salazar A, Diego F L. Fine-grained recognition of wildlife in camera trap images with vision transformers. Remote Sensing in Ecology and Conservation, Dean&Francis Yunke Wang 2022, 8(2): 123-137.

[4] Shorten C, Khoshgoftaar T M. A survey on image data augmentation for deep learning. Journal of Big Data, 2019, 6(1): 1-48.

[5] Howard J, Ruder S. Universal language model fine-tuning for text classification. Annual Meeting of the Association for Computational Linguistics, 2018: 328-339.

[6] Krishnamoorthi R. Quantizing deep convolutional networks for efficient inference: A whitepaper. arXiv preprint arXiv:1806.08342, 2018.

[7] Yang X, Yu S, Xu W. Enhanced convolutional neural networks for improved image classification. arXiv preprint arXiv:2502.00663, 2025.

[8] Huang Y. Using neural networks to build an efficient classification model for classifying images in the CIFAR-10 dataset. International Conference on Neural Networks, 2024.

[9] Zhu L, Yang D, Xu X. A transformer-based model for wildlife species identification using camera trap images. IEEE International Conference on Image Processing, 2022: 2087- 2091.

[10] Zhang W, et al. CoTr: Efficiently bridging CNN and transformer for medical image segmentation. International Conference on Medical Image Computing and Computer- Assisted Intervention, 2021: 213-223.

[11] Bansal A, Khurana G. Advancing image classification performance: A comprehensive study of modern deep learning architectures on CIFAR-10. Global Journal of Computer Science and Technology, 2025, 25(F1): 21-27.

[12] Pant Y, Shah G, Ojha R. Comparison of CNN architectures for image classification using CIFAR-10 dataset. International Journal on Engineering Technology, 2023, 1(1): 37-52.

[13] Ghafouri S. Enhancing image classification accuracy using convolutional neural network on CIFAR-10 dataset. University of Victoria Technical Report, 2024.

Downloads

Published

2026-02-28