paper-with-me

홈 › Papers

Source framing triggers systematic evaluation bias in Large Language Models

2025-05-14 · Federico Germani, Giovanni Spitale

Large Language Models (LLMs) are increasingly used not only to generate text but also to evaluate it, raising urgent questions about whether their judgments are consistent, unbiased, and robust to framing effects. In this study, we systematically examine inter- and intra-model agreement across four state-of-the-art LLMs (OpenAI o3-mini, Deepseek Reasoner, xAI Grok 2, and Mistral) tasked with evaluating 4,800 narrative statements on 24 different topics of social, political, and public health relevance, for a total of 192,000 assessments. We manipulate the disclosed source of each statement to assess how attribution to either another LLM or a human author of specified nationality affects evaluation outcomes. We find that, in the blind condition, different LLMs display a remarkably high degree of inter- and intra-model agreement across topics. However, this alignment breaks down when source framing is introduced. Here we show that attributing statements to Chinese individuals systematically lowers agreement scores across all models, and in particular for Deepseek Reasoner. Our findings reveal that framing effects can deeply affect text evaluation, with significant implications for the integrity, neutrality, and fairness of LLM-mediated information systems.

📄 PDF Abstract BibTeX arXiv:2505.13488

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

When Wording Steers the Evaluation: Framing Bias in LLM judges

2026-01-20 · Yerin Hwang, Dongryeol Lee, Taegwan Kang, Minwoo Lee 외 arxiv

Large language models (LLMs) are known to produce varying responses depending on prompt phrasing, indicating that subtle guidance in phrasing can steer their answers. However, the impact of this framing bias on LLM-based…

DeFrame: Debiasing Large Language Models Against Framing Effects

2026-02-04 · Kahee Lim, Soyeon Kim, Steven Euijong Whang arxiv

As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has become crucial. Despite many efforts, an ongoing challenge is hidden bias: LLMs …

Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models

2026-01-15 · Abhinaba Basu, Pavan Chakraborty arxiv

A model that avoids stereotypes in a lab benchmark may not avoid them in deployment. We show that measured bias shifts dramatically when prompts mention different places, times, or audiences -- no adversarial prompting r…

FrameRef: A Framing Dataset and Simulation Testbed for Modeling Bounded Rational Information Health

2026-02-17 · Victor De Lima, Jiqun Liu, Grace Hui Yang arxiv

Information ecosystems increasingly shape how people internalize exposure to adverse digital experiences, raising concerns about the long-term consequences for information health. In modern search and recommendation syst…

Recommendation Systems

BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts

2026-01-11 · William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault 외 arxiv

Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisi…