Editable Fairness: Fine-Grained Bias Mitigation in Language Models

Chen, Ruizhe; Li, Yichen; Yang, Jianfei; Zhou, Joey Tianyi; Liu, Zuozhu

Computer Science > Computation and Language

arXiv:2408.11843 (cs)

[Submitted on 7 Aug 2024]

Title:Editable Fairness: Fine-Grained Bias Mitigation in Language Models

Authors:Ruizhe Chen, Yichen Li, Jianfei Yang, Joey Tianyi Zhou, Zuozhu Liu

View PDF HTML (experimental)

Abstract:Generating fair and accurate predictions plays a pivotal role in deploying large language models (LLMs) in the real world. However, existing debiasing methods inevitably generate unfair or incorrect predictions as they are designed and evaluated to achieve parity across different social groups but leave aside individual commonsense facts, resulting in modified knowledge that elicits unreasonable or undesired predictions. In this paper, we first establish a new bias mitigation benchmark, BiaScope, which systematically assesses performance by leveraging newly constructed datasets and metrics on knowledge retention and generalization. Then, we propose a novel debiasing approach, Fairness Stamp (FAST), which enables fine-grained calibration of individual social biases. FAST identifies the decisive layer responsible for storing social biases and then calibrates its outputs by integrating a small modular network, considering both bias mitigation and knowledge-preserving demands. Comprehensive experiments demonstrate that FAST surpasses state-of-the-art baselines with superior debiasing performance while not compromising the overall model capability for knowledge retention and downstream predictions. This highlights the potential of fine-grained debiasing strategies to achieve fairness in LLMs. Code will be publicly available.

Comments:	arXiv admin note: substantial text overlap with arXiv:2405.09341
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2408.11843 [cs.CL]
	(or arXiv:2408.11843v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2408.11843

Submission history

From: Ruizhe Chen [view email]
[v1] Wed, 7 Aug 2024 17:14:58 UTC (6,079 KB)

Computer Science > Computation and Language

Title:Editable Fairness: Fine-Grained Bias Mitigation in Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Editable Fairness: Fine-Grained Bias Mitigation in Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators