Core Applications and Techniques of RAG for Low-Resource Languages
DOI:
https://doi.org/10.61173/745jfc20Keywords:
Low-resource Languages, Retrieval-Augmented Generation (RAG), Cross-lingual Information Retrieval, Digital InclusionAbstract
Low-resource languages (LRLs) are spoken by billions of people globally, yet they continue to face limited access to reliable AI-driven tools, largely due to a lack of sufficient digital text resources. This review examines recent progress in Retrieval-Augmented Generation (RAG) that addresses the specific difficulties encountered in processing LRLs. By analyzing key studies, a range of optimization strategies are identified and analyzed across the RAG workflow—including data handling, retrieval techniques, and generation refinements. The discussion emphasizes how these methods tackle fundamental problems such as scarce data, weak model performance, and poor cultural alignment. The role of cross-lingual retrieval and knowledge distillation is also explored as a way to make AI systems more accessible and useful for speakers of low-resource languages. In addition to outlining RAG’s potential for reducing language-based digital inequality, this survey notes remaining obstacles and suggests productive avenues for further research. It is hoped that this structured summary of methods and use cases will support future efforts toward inclusive AI and the protection of linguistic diversity.
References
[1] Yu, P., Fei, H., & Li, P. (2021). Cross-lingual language model pretraining for retrieval. In Proceedings of the Web Conference 2021 (WWW ’21). ACM. https://doi. org/10.1145/3442381.3449830
[2] Jiang, Z., El-Jaroudi, A., Hartmann, W., Karakos, D., & Zhao, L. (2020). Cross-lingual information retrieval with BERT. arXiv preprint arXiv:2004.13005. https://arxiv.org/abs/2004.13005
[3] Huang, Z., Yu, P., & Allan, J. (2023). Improving crosslingual information retrieval on low-resource languages via optimal transport distillation. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining (WSDM ’23). ACM. https://doi.org/10.1145/3539597.3570468
[4] Alshammary, M., Uddin, M. N., & Khan, L. (2024). RFPG: Question-Answering from Low-Resource Language (Arabic) Texts using Factually Aware RAG. In Proceedings of the 2024 IEEE 10th International Conference on Collaboration and Internet Computing (CIC) (pp. 107-116). IEEE. https://doi. org/10.1109/CIC62241.2024.00023 Dean&Francis ISSN 2959-6157
[5] Dutta, B., Ranjan, R., Jain, A., Singh, R., & Vatsa, M. (2025). Can RAG-driven enhancements amplify audio LLMs for low-resource languages? In Proceedings of the 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE. https://doi.org/10.1109/ ICASSP49660.2025.10889964
[6] Li, Z., & Ke, Z. (2025). Cross-Modal Augmentation for Low-Resource Language Understanding and Generation. In Proceedings of the 1st Workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR 2025) (pp. 90–99). Association for Computational Linguistics. https://doi. org/10.18653/v1/2025.magmar-1.9
[7] Seo, M., Baek, J., Thorne, J., & Hwang, S. J. (2024). Retrieval-augmented data augmentation for low-resource domain tasks. arXiv preprint arXiv:2402.13482. https://arxiv. org/abs/2402.13482
[8] Nie, E., Liang, S., Schmid, H., & Schütze, H. (2023). Cross-lingual retrieval augmented prompt for low-resource languages. arXiv preprint arXiv:2212.09651. https://arxiv.org/ abs/2212.09651
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
