The Development and Future Outlook of CLIP and its Derivative Methods

Authors

  • Jiacheng Shi

DOI:

https://doi.org/10.61173/0p6jpg91

Keywords:

Contrastive Language–Image Pre-training, Contrastive learning, Vision-Language Models

Abstract

With the rapid growth of multimodal learning, VisionLanguage Models (VLMs) have become a cutting-edge direction in artificial intelligence. Among them, the Contrastive Language–Image Pre-training (CLIP) model, based on large-scale contrastive learning, has demonstrated powerful capabilities in zero-shot transfer and crossmodal retrieval. However, CLIP’s weakly supervised training paradigm shows clear shortcomings when dealing with compositional reasoning. Therefore, this survey systematically reviews and analyzes representative methods proposed in recent years to address CLIP’s compositional reasoning limitations, including Self-supervision meets Language-Image Pre-training (SLIP), Language augmented CLIP (LaCLIP), TripletCLIP, Synthetic Perturbations for Advancing Robust Compositional Learning (SPARCL), Compositionally-aware Learning in CLIP (CLIC), and Training-Time Negation Data Generation for Negation Awareness of CLIP (TNG-CLIP). We introduce the principles and characteristics of these methods, followed by a comparative analysis of their performance on different benchmarks and how they mitigate deficiencies. Through this overview of CLIP and its derivative methods, we hope future research will focus on integrating their strengths, while also developing more efficient data synthesis techniques and more comprehensive evaluation benchmarks.

References

[1] Tan H, & Bansal M. Vokenization: Improving language understanding with contextualized, visual-grounded supervision. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), 2020, 2066–2080.

[2] Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, ... & Sutskever I. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML 2021), 2021, 13921–13935.

[3] Mu N, Kirillov A, Wagner D, & Xie S. SLIP: Selfsupervision Meets Language-Image Pre-training. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds) Computer Vision (ECCV 2022), 2022, 13686.

[4] Fan L, Krishnan D, Isola P, Katabi D, & Tian Y. Improving CLIP training with language rewrites. Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023.

[5] Patel M, Kusumba A, Cheng S, Kim C, Gokhale T, Baral C & Yang Y. TripletCLIP: Improving compositional reasoning of CLIP via synthetic vision-language negatives. Part of Advances in Neural Information Processing Systems 37 (NeurIPS 2024), 2024.

[6] Cai Y, Thomason J, & Rostami M. TNG-CLIP: Training-time negation data generation for negation awareness of CLIP. arXiv preprint arXiv:2505.18434, 2025.

[7] Li H, & Li B. Enhancing vision-language compositional understanding with Multimodal Synthetic Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2025), 2025, 24849-24861.

[8] Peleg A, Singh ND, & Hein M. Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning. arXiv preprint arXiv:2505.24424, 2025.

[9] Changpinyo S, Sharma P, Ding N, & Soricut R. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2021), 2021, 3317–3327.

[10] Desai K, Kaul G, Aysola Z, & Johnson J. RedCaps: Webcurated image-text data created by the people, for the people. In Proceedings of the 35th Conference on Neural Information Dean&Francis ISSN 2959-6157 Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks, 2021.

[11] Thrush T, Jiang R, Bartolo M, Singh A, Williams A, Kiela D, & Ross C. Winoground: Probing vision and language models for compositional understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), 2022, 5238-5248

[12] Zerroug A, Vaishnav M, Colin J, Musslick S, & Serre T. A benchmark for compositional visual reasoning. Part of Advances in Neural Information Processing Systems 35 (NeurIPS 2022) Datasets and Benchmarks Track, 2022.

[13] Lin TY, Maire M, Belongie S, Hays J, Perona P, ... & Zitnick P. Microsoft COCO: Common Objects in Context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds) Computer Vision – ECCV 2014. ECCV 2014. Lecture Notes in Computer Science, vol 8693. Springer, Cham. 2014.

[14] Plummer BA, Wang L, Cervantes CM, Caicedo JC, Hockenmaier J, & Lazebnik S. Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2015), 2015, 4116–4124.

[15] Deng J, Dong W, Socher R, Li LJ, Li K, Li FF. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2009), 2009, 248–255.

Downloads

Published

2025-12-19