paper-with-me

홈 › Papers

Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models

2024-12-02 · Schrasing Tong, Eliott Zemour, Rawisara Lohanimit, Lalana Kagal

Although large language models (LLMs) have demonstrated their effectiveness in a wide range of applications, they have also been observed to perpetuate unwanted biases present in the training data, potentially leading to harm for marginalized communities. In this paper, we mitigate bias by leveraging small biased and anti-biased expert models to obtain a debiasing signal that will be added to the LLM output at decoding-time. This approach combines resource efficiency with interpretability and can be optimized for mitigating specific types of bias, depending on the target use case. Experiments on mitigating gender, race, and religion biases show a reduction in bias on several local and global bias metrics while preserving language model performance.

📄 PDF Abstract BibTeX arXiv:2412.01711

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs

2026-01-30 · Afrozah Nadeem, Agrima Seth, Mehwish Nasim, Usman Naseem arxiv

Large Language Models (LLMs) increasingly shape global discourse, making fairness and ideological neutrality essential for responsible AI deployment. Despite growing attention to political bias in LLMs, prior work largel…

CORGI-PM: A Chinese Corpus For Gender Bias Probing and Mitigation

2023-01-01 · Ge Zhang, Yizhi Li, Yaoyao Wu, Linyuan Zhang 외

As natural language processing (NLP) for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques such as large-scale language models suffer from data inadequacy and biased corpus, …

Sentence

MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion

2025-07-03 · Xin Guan, PeiHsin Lin, Zekun Wu, Ze Wang 외 arxiv

Multiperspective Fusion (MPF) is a novel posttraining alignment framework for large language models (LLMs) developed in response to the growing need for easy bias mitigation. Built on top of the SAGED pipeline, an automa…

Prompt Engineering

From Measurement to Mitigation: Exploring the Transferability of Debiasing Approaches to Gender Bias in Maltese Language Models

2025-07-03 · Melanie Galea, Claudia Borg arxiv

The advancement of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), enabling performance across diverse tasks with little task-specific training. However, LLMs remain susceptible to social …

Data Augmentation

Gender Bias Mitigation for Bangla Classification Tasks

2024-11-16 · Sajib Kumar Saha Joy, Arman Hassan Mahy, Meherin Sultana, Azizah Mamun Abha 외

In this study, we investigate gender bias in Bangla pretrained language models, a largely under explored area in low-resource languages. To assess this bias, we applied gender-name swapping techniques to existing dataset…

ClassificationHate Speech DetectionSarcasm DetectionSentiment Analysis