Echoes of Agreement: Argument Driven Opinion Shifts in Large Language Models
There have been numerous studies evaluating bias of LLMs towards political topics. However, how positions towards these topics in model outputs are highly sensitive to the prompt. What happens when the prompt itself is suggestive of certain arguments towards those positions remains underexplored. This is crucial for understanding how robust these bias evaluations are and for understanding model behaviour, as these models frequently interact with opinionated text. To that end, we conduct experiments for political bias evaluation in presence of supporting and refuting arguments. Our experiments show that such arguments substantially alter model responses towards the direction of the provided argument in both single-turn and multi-turn settings. Moreover, we find that the strength of these arguments influences the directional agreement rate of model responses. These effects point to a sycophantic tendency in LLMs adapting their stance to align with the presented arguments which has downstream implications for measuring political bias and developing effective mitigation strategies.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Annotating Arguments in a Corpus of Opinion Articles
Interest in argument mining has resulted in an increasing number of argument annotated corpora. However, most focus on English texts with explicit argumentative discourse markers, such as persuasive essays or legal docum…
Argument MiningArticlesAnnotating argumentation in Swedish social media
This paper presents a small study of annotating argumentation in Swedish social media. Annotators were asked to annotate spans of argumentation in 9 threads from two discussion forums. At the post level, Cohen’s k and Kr…
Agreement Prediction of Arguments in Cyber Argumentation for Detecting Stance Polarity and Intensity
In online debates, users express different levels of agreement/disagreement with one another{'}s arguments and ideas. Often levels of agreement/disagreement are implicit in the text, and must be predicted to analyze coll…
regressionStance DetectionDisentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
Large Language Models are increasingly used to simulate human opinion dynamics, yet the effect of genuine interaction is often obscured by systematic biases. We develop a Bayesian framework to disentangle and quantify th…
A Corpus for Research on Deliberation and Debate
Deliberative, argumentative discourse is an important component of opinion formation, belief revision, and knowledge discovery; it is a cornerstone of modern civil society. Argumentation is productively studied in branch…