Attention-Centric YOLOv12 for Real-Time Fine-Grained Waste Detection in the TACO Dataset

Authors

  • Hongye Wu

DOI:

https://doi.org/10.61173/aexy0p38

Keywords:

YOLOv12, Waste detection, TACO dataset, Attention mechanism, Real-time inference

Abstract

Efficient waste detection is crucial for environmental sustainability, yet existing models struggle with finegrained objects in complex backgrounds, such as those in the TACO dataset. This paper proposes an attention-centric approach using YOLOv12n to balance detection accuracy and real-time performance. Experimental results on the TACO dataset demonstrate that the proposed YOLOv12n achieves a mean Average Precision (mAP50 ) of 0.376 with an end-to-end inference speed of 94.33 FPS on an NVIDIA RTX 5060 GPU. Ablation studies reveal that removing the Area-Attention ( A2 ) module leads to a significant performance drop, with mAP50 plummeting from 0.376 to 0.148. Furthermore, compared to the YOLOv8n model with a plug-in CBAM module (46.18 FPS, 0.310 mAP), the native attention-centric architecture of YOLOv12n provides a significant increase in inference speed and superior feature localization. This research confirms that a native attention-based design is more effective for realtime fine-grained waste detection than traditional modular additions.

References

[1] P. F. Proença and P. Simões, “TACO: A Trash Annotations in Context Dataset for Litter Detection,” arXiv preprint arXiv:2003.06975, 2020.

[2] G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics YOLOv8,” 2023. [Online]. Available: https://github.com/ultralytics/ ultralytics.

[3] S. Woo, J. Park, J. Y. Lee, and I. S. Kweon, “CBAM: Convolutional Block Attention Module,” in Proc. Eur. Conf. Comput. Vis. (ECCV), 2018, pp. 3-19.

[4] W. Wang et al., “YOLOv12: Attention-Centric Real-Time Object Detectors,” arXiv preprint arXiv:2502.12524, 2025.)

[5] A. Vaswani et al., “Attention is All You Need,” in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017, pp. 5998-6008.

[6] M. Yang and G. Thung, “Classification of Trash for Recyclability Status,” CS229 Project Report, Stanford University, 2016.

[7] C. Y. Wang, I. H. Yeh, and H. Y. M. Liao, “YOLOv9: Learning What You Want to Learn Through Programmable Gradient Information,” arXiv preprint arXiv:2402.13616, 2024.

[8] A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.

[9] Z. Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 10012-10022.

[10] T. Y. Lin et al., “Focal Loss for Dense Object Detection,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017, pp. 2980- 2988.

[11] S. Hu, J. Zhang, and J. Lu, “A Survey on Object Detection for Intelligent Waste Management,” IEEE Access, vol. 10, pp. 12345-12360, 2022.

[12] H. Wang et al., “YOLOv10: Real-Time End-to-End Object Detection,” arXiv preprint arXiv:2405.14458, 2024.

[13] M. Sandler et al., “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 4510-4520.

[14] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 770-778.

[15] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2016, pp. 779-788.

Downloads

Published

2026-04-24