paper-with-me

Papers

Don't Change My View: Ideological Bias Auditing in Large Language Models

2025-09-16 · Paul Kröger, Emilio Barkett arxiv

As large language models (LLMs) become increasingly embedded in products used by millions, their outputs may influence individual beliefs and, cumulatively, shape public opinion. If the behavior of LLMs can be intentionally steered toward specific ideological positions, such as political or religious views, then those who control these systems could gain disproportionate influence over public discourse. Although it remains an open question whether LLMs can reliably be guided toward coherent ideological stances and whether such steering can be effectively prevented, a crucial first step is to develop methods for detecting when such steering attempts occur. In this work, we adapt a previously proposed statistical method to the new context of ideological bias auditing. Our approach carries over the model-agnostic design of the original framework, which does not require access to the internals of the language model. Instead, it identifies potential ideological steering by analyzing distributional shifts in model outputs across prompts that are thematically related to a chosen topic. This design makes the method particularly suitable for auditing proprietary black-box systems. We validate our approach through a series of experiments, demonstrating its practical applicability and its potential to support independent post hoc audits of LLM behavior.

📄 PDF Abstract BibTeX arXiv:2509.12652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias

2026-01-08 · Adib Sakhawat, Tahsin Islam, Takia Farhin, Syed Rifat Raiyan 외 arxiv

As large language models (LLMs) are increasingly deployed, understanding how they express political positioning is important for evaluating alignment and downstream effects. We audit 26 contemporary LLMs using three poli…

Uncovering Political Bias in Large Language Models using Parliamentary Voting Records

2026-01-13 · Jieying Chen, Karen de Jong, Andreas Poole, Jan Burakowski 외 arxiv

As large language models (LLMs) become deeply embedded in digital platforms and decision-making systems, concerns about their political biases have grown. While substantial work has examined social biases such as gender …

GermanPartiesQA: Benchmarking Commercial Large Language Models for Political Bias and Sycophancy

2024-07-25 · Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci

LLMs are changing the way humans create and interact with content, potentially affecting citizens' political opinions and voting decisions. As LLMs increasingly shape our digital information ecosystems, auditing to evalu…

Benchmarking

Unsupervised Detection of Contextualized Embedding Bias with Application to Ideology

2022-12-14 · Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze

We propose a fully unsupervised method to detect bias in contextualized embeddings. The method leverages the assortative information latently encoded by social networks and combines orthogonality regularization, structur…

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

2025-08-31 · Sanjeeevan Selvaganapathy, Mehwish Nasim arxiv

We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) cou…

Hate Speech Detection