Bias Mitigation Techniques in Large Language Models
DOI:
https://doi.org/10.61173/4yvqsc19Keywords:
Large language models (LLMs), Prejudice and fairness, Bias mitigation techniquesAbstract
Large scale language models (LLMs) demonstrate outstanding performance and enormous potential for development, and are widely applied in people’s real-life situations. However, social bias can be learned by LLM in unprocessed training data and transmitted to downstream tasks, resulting in adverse social effects and potential harm. In this article, we present a survey of bias and fairness research on Large Language Models (LLMs), categorizing the metrics and datasets used for bias assessment. Based on the elements used by the metrics in the model, they are refined into embeddings, probabilities, and generated text. The dataset is then divided into counterfactual inputs or prompts based on its structure. Afterwards, this article conducts research and organization on bias mitigation techniques based on different intervention stages: preprocessing (modifying model inputs), in-training (modifying optimization processes), intra-processing (modifying inference behavior), and post-processing (modifying model outputs). Finally, this study aims to explore in depth the key challenges that affect the fair development of large language models, and to look forward to their future evolution paths.
References
[1] Baeza-Yates R. Data and algorithmic bias in the web// Proceedings of the 8th ACM Conference on Web Science. 2016: 1-1.
[2] Bansal R. A survey on bias and fairness in natural language processing. arXiv preprint arXiv:2204.09591, 2022.
[3] Mehrabi N, Morstatter F, Saxena N, et al. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR), 2021, 54(6): 1-35.
[4] Gohar U, Cheng L. A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges. arXiv preprint arXiv:2305.06969, 2023.
[5] Ferrara E. Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies. Sci, 2024, 6(1): 3.
[6] Navigli R, Conia S, Ross B. Biases in large language models: origins, inventory, and discussion. ACM Journal of Data and Information Quality, 2023, 15(2): 1-21.
[7] Wang P, Li L, Chen L, et al. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023.
[8] Ghanbarzadeh S, Huang Y, Palangi H, et al. Gender-tuning: Empowering fine-tuning for debiasing pre-trained language models. arXiv preprint arXiv:2307.10522, 2023.
[9] Zayed A, Parthasarathi P, Mordido G, et al. Deep learning on a healthy data diet: Finding important examples for fairness// Proceedings of the AAAI Conference on Artificial Intelligence. 2023, 37(12): 14593-14601.
[10] Yu L, Mao Y, Wu J, et al. Mixup-based unified framework to overcome gender bias resurgence//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2023: 1755-1759.
[11] Thakur H, Jain A, Vaddamanu P, et al. Language models get a gender makeover: Mitigating gender bias with few-shot data interventions. arXiv preprint arXiv:2306.04597, 2023.
[12] Han X, Baldwin T, Cohn T. Balancing out bias: Achieving fairness through balanced training. arXiv preprint arXiv:2109.08253, 2021.
[13] Omrani A, Salkhordeh_Ziabari A, Yu C, et al. Socialgroup-agnostic bias mitigation via the stereotype content model. Association for Computational Linguistics, 2023.
[14] Hauzenberger L, Masoudian S, Kumar D, et al. Modular and on-demand bias mitigation with attribute-removal subnetworks. arXiv preprint arXiv:2205.15171, 2022
[15] Amrhein C, Schottmann F, Sennrich R, et al. Exploiting biased models to de-bias text: A gender-fair rewriting model. arXiv preprint arXiv:2305.11140, 2023.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
