paper-with-me

홈 › Papers

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

2026-04-02 · Lucas Resck, Isabelle Augenstein, Anna Korhonen arxiv

Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptation is generally desirable, it becomes a critical failure when a user's identity is explicitly defined. For instance, given a fixed British persona and an ambiguous everyday knowledge query about literature, the prompt's language frequently overwrites the system persona -- yielding Shakespeare in English but Cervantes in Spanish. To robustly quantify this Cross-lingual Cultural Inconsistency, we introduce Singleton Fleiss's $κ_S$, a metric mathematically resilient to hallucinations. For mitigation, we propose Cross-lingual Cultural Consistent Preference Optimisation (C-3PO), a consensus-driven alignment framework. C-3PO achieves up to a 0.13-point absolute increase in $κ_S$ over unaligned models, consistently outperforming strong prompting and representation steering baselines whilst preserving explicit user identities, cultural neutrality and intrinsic cultural knowledge. Empirical evaluations demonstrate this inconsistency disproportionately affects lower-resource languages like Indonesian and Persian. Finally, early decoding of intermediate layers reveals that MLLMs implicitly personalise outputs towards the prompt language's stereotypical culture as forward-pass representations stabilise.

📄 PDF Abstract BibTeX arXiv:2605.12515

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning

2025-09-25 · Sualeha Farid, Jayden Lin, Zean Chen, Shivani Kumar 외 arxiv

Large Language Models (LLMs) are increasingly deployed in multilingual and multicultural environments where moral reasoning is essential for generating ethically appropriate responses. Yet, the dominant pretraining of LL…

Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering

2025-10-13 · Lorena Calvo-Bartolomé, Valérie Aldana, Karla Cantarero, Alonso Madroñal de Mesa 외 arxiv

Multilingual question answering (QA) systems must ensure factual consistency across languages, especially for objective queries such as What is jaundice?, while also accounting for cultural variation in subjective respon…

Question Answering

Evaluating Knowledge-based Cross-lingual Inconsistency in Large Language Models

2024-07-01 · Xiaolin Xing, Zhiwei He, Haoyu Xu, Xing Wang 외

This paper investigates the cross-lingual inconsistencies observed in Large Language Models (LLMs), such as ChatGPT, Llama, and Baichuan, which have shown exceptional performance in various Natural Language Processing (N…

Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation

2025-10-10 · Muhammad Ali Shafique, Kanwal Mehreen, Muhammad Arham, Maaz Amjad 외 arxiv

Developing a high-performing large language models (LLMs) for low-resource languages such as Urdu, present several challenges. These challenges include the scarcity of high-quality datasets, multilingual inconsistencies,…

MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs

2025-05-27 · Raoyuan Zhao, Beiduo Chen, Barbara Plank, Michael A. Hedderich

Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness, often resulting in biased outputs. Howev…

SpecificityText GenerationTranslation