paper-with-me

Papers

GermanPartiesQA: Benchmarking Commercial Large Language Models for Political Bias and Sycophancy

2024-07-25 · Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci

LLMs are changing the way humans create and interact with content, potentially affecting citizens' political opinions and voting decisions. As LLMs increasingly shape our digital information ecosystems, auditing to evaluate biases, sycophancy, or steerability has emerged as an active field of research. In this paper, we evaluate and compare the alignment of six LLMs by OpenAI, Anthropic, and Cohere with German party positions and evaluate sycophancy based on a prompt experiment. We contribute to evaluating political bias and sycophancy in multi-party systems across major commercial LLMs. First, we develop the benchmark dataset GermanPartiesQA based on the Voting Advice Application Wahl-o-Mat covering 10 state and 1 national elections between 2021 and 2023. In our study, we find a left-green tendency across all examined LLMs. We then conduct our prompt experiment for which we use the benchmark and sociodemographic data of leading German parliamentarians to evaluate changes in LLMs responses. To differentiate between sycophancy and steerabilty, we use 'I am [politician X], ...' and 'You are [politician X], ...' prompts. Against our expectations, we do not observe notable differences between prompting 'I am' and 'You are'. While our findings underscore that LLM responses can be ideologically steered with political personas, they suggest that observed changes in LLM outputs could be better described as personalization to the given context rather than sycophancy.

📄 PDF Abstract BibTeX arXiv:2407.18008

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

Aligning Large Language Models with Diverse Political Viewpoints

2024-06-20 · Dominik Stammbach, Philine Widmer, Eunjung Cho, Caglar Gulcehre 외

Large language models such as ChatGPT exhibit striking political biases. If users query them about political information, they often take a normative stance. To overcome this, we align LLMs with diverse political viewpoi…

Benchmarking Gender and Political Bias in Large Language Models

2025-09-07 · Jinrui Yang, Xudong Han, Timothy Baldwin arxiv

We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-call vote outcomes and includes rich demogr…

Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models

2025-03-20 · Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz 외

Prompt-based language models like GPT4 and LLaMa have been used for a wide variety of use cases such as simulating agents, searching for information, or for content analysis. For all of these applications and others, pol…

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

2026-08-18 · Yijie Xu, Chao Wang, Hui Xiong arxiv

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarci…

Unsupervised Domain AdaptationStyle Transfer

AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI

2025-01-09 · Pavlos Sermpezis, Stelios Karamanidis, Eva Paraschou, Ilias Dimitriadis 외

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality,…

Benchmarkingnamed-entity-recognitionNamed Entity RecognitionSentiment Analysis+2