paper-with-me

홈 › Papers

Mitigating Cultural Bias in LLMs via Multi-Agent Cultural Debate

2026-01-17 · Qian Tan, Lei Jiang, Yuting Zeng, Shuoyang Ding, Xiaohua Xu arxiv

Large language models (LLMs) exhibit systematic Western-centric bias, yet whether prompting in non-Western languages (e.g., Chinese) can mitigate this remains understudied. Answering this question requires rigorous evaluation and effective mitigation, but existing approaches fall short on both fronts: evaluation methods force outputs into predefined cultural categories without a neutral option, while mitigation relies on expensive multi-cultural corpora or agent frameworks that use functional roles (e.g., Planner--Critique) lacking explicit cultural representation. To address these gaps, we introduce CEBiasBench, a Chinese--English bilingual benchmark, and Multi-Agent Vote (MAV), which enables explicit ``no bias'' judgments. Using this framework, we find that Chinese prompting merely shifts bias toward East Asian perspectives rather than eliminating it. To mitigate such persistent bias, we propose Multi-Agent Cultural Debate (MACD), a training-free framework that assigns agents distinct cultural personas and orchestrates deliberation via a "Seeking Common Ground while Reserving Differences" strategy. Experiments demonstrate that MACD achieves 57.6% average No Bias Rate evaluated by LLM-as-judge and 86.0% evaluated by MAV (vs. 47.6% and 69.0% baseline using GPT-4o as backbone) on CEBiasBench and generalizes to the Arabic CAMeL benchmark, confirming that explicit cultural representation in agent frameworks is essential for cross-cultural fairness.

📄 PDF Abstract BibTeX arXiv:2601.12091

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation

2025-09-25 · Yixin Wan, Xingrun Chen, Kai-Wei Chang arxiv

Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. However, recent research raised concerns about culture-related fairness issues in…

WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

2025-05-14 · Abdullah Mushtaq, Imran Taj, Rafay Naeem, Ibrahim Ghaznavi 외

Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenization and limiting their ability to reflect …

Benchmarking

MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints

2025-08-07 · Zhong Ken Hew, Jia Xin Low, Sze Jue Yang, Chee Seng Chan arxiv

Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cult…

Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems

2025-03-05 · Mahfuz Ahmed Anik, Abdur Rahman, Azmine Toushik Wasi, Md Manjurul Ahsan

Language is a cornerstone of cultural identity, yet globalization and the dominance of major languages have placed nearly 3,000 languages at risk of extinction. Existing AI-driven translation models prioritize efficiency…

DiversityTranslation

Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens

2025-01-02 · Abdullah Mushtaq, Muhammad Rafay Naeem, Muhammad Imran Taj, Ibrahim Ghaznavi 외

As large language models (LLMs) like GPT-4 and Llama 3 become integral to educational contexts, concerns are mounting over the cultural biases, power imbalances, and ethical limitations embedded within these technologies…

Sentiment Analysis