paper-with-me

홈 › Papers

Warning labels shift perceptions of sycophantic AI, but not its influence

2026-06-19 · Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe, Desmong Ong, Dan Jurafsky, Diyi Yang arxiv

Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which has received regulatory attention, is to warn users about potentially harmful AI behaviors such as sycophancy. In a preregistered experiment in which participants (N = 2,610) discussed real interpersonal conflicts with an AI system, we test whether warning labels mitigate sycophancy's influence. We find that a basic AI disclosure (`This chatbot is AI'') has no detectable effect. Labeling the system as sycophantic (`...may agree with you and validate you even when you are wrong...'') does shift users' perceptions, reducing perceived objectivity and trust, but it does not reliably reduce sycophancy's influence on users' self-perceived rightness or their willingness to repair the conflict. Our results reveal a gap between AI perception and AI influence: by shifting perception without reducing influence, warning-based interventions may offer a false sense of protection. Addressing the harms of sycophancy will therefore require understanding the specific mechanisms through which it shapes judgment, and improving model behavior itself.

📄 PDF Abstract BibTeX arXiv:2606.21317

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

2026-07-28 · Meryl Ye, Robert Kraut, Steve Rathje arxiv

AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindnes…

Labeling Synthetic Content: User Perceptions of Warning Label Designs for AI-generated Content on Social Media

2025-02-14 · Dilrukshi Gamage, Dilki Sewwandi, Min Zhang, Arosha Bandara

In this research, we explored the efficacy of various warning label designs for AI-generated content on social media platforms e.g., deepfakes. We devised and assessed ten distinct label design samples that varied across…

Face Swapping

Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model

2024-12-03 · María Victoria Carro

Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factual…

Language ModelingLanguage ModellingLarge Language ModelMisinformation

More Human or More AI? Visualizing Human-AI Collaboration Disclosures in Journalistic News Production

2026-01-16 · Amber Kusters, Pooja Prajod, Pablo Cesar, Abdallah El Ali arxiv

Within journalistic editorial processes, disclosing AI usage is currently limited to simplistic labels, which misses the nuance of how humans and AI collaborated on a news article. Through co-design sessions (N=10), we e…

Sycophantic AI makes human interaction feel more effortful and less satisfying over time

2026-05-08 · Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng, Cinoo Lee 외 arxiv

Millions of people now turn to artificial intelligence (AI) systems for personal advice, guidance, and support. Such systems can be sycophantic, frequently affirming users' views and beliefs. Across five preregistered st…