paper-with-me

홈 › Papers

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

2026-03-21 · Jocelyn Shen, Amina Luvsanchultem, Jessica Kim, Kynnedy Smith, Valdemar Danry, Kantwon Rogers, Hae Won Park, Maarten Sap, Cynthia Breazeal arxiv

As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned with their own interests. While existing NLP research has benchmarked manipulation detection, these efforts often rely on simulated debates and remain fundamentally decoupled from actual human belief shifts in real-world scenarios. We introduce PUPPET, a theoretical taxonomy and resource that bridges this gap by focusing on the moral direction of hidden incentives in everyday, advice-giving contexts. We provide an evaluation dataset of N=1,035 human-LLM interactions, where we measure users' belief shifts. Our analysis reveals a critical disconnect in current safety paradigms: while models can be trained to detect manipulative strategies, they do not correlate with the magnitude of resulting belief change. As such, we define the task of belief shift prediction and show that while state-of-the-art LLMs achieve moderate correlation (r=0.3-0.5), they exhibit systematic directional biases, with some models over-predicting and others under-predicting the magnitude of human belief change. This work establishes a theoretically grounded and behaviorally validated foundation for AI social safety efforts by studying incentive-driven manipulation in LLMs during everyday, practical user queries.

📄 PDF Abstract BibTeX arXiv:2603.20907

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics

2024-08-08 · Ruining Li, Chuanxia Zheng, Christian Rupprecht, Andrea Vedaldi

We present Puppet-Master, an interactive video generative model that can serve as a motion prior for part-level dynamics. At test time, given a single image and a sparse set of motion trajectories (i.e., drags), Puppet-M…

Video Generation

Detecting Sockpuppets in Deceptive Opinion Spam

2017-03-09 · Marjan Hosseinia, Arjun Mukherjee

This paper explores the problem of sockpuppet detection in deceptive opinion spam using authorship attribution and verification approaches. Two methods are explored. The first is a feature subsampling scheme that uses th…

Authorship AttributionDiversity

Invisible Strings: Deriving Puppetry Principles and their Hidden Connections to Robot Behavior Design

2026-07-03 · Claire Lewis, Sawyer Collins, Alyssa Hanson, Johanna Smith 외 arxiv

When designing robots' nonverbal behaviors, many researchers have turned to arts-based insights, such as Disney's Animation Principles. Yet, while these principles bear key insights into the design of like-life character…

The Stitched Puppet: A Graphical Model of 3D Human Shape and Pose

2015-06-01 · CVPR 2015 6 · Silvia Zuffi, Michael J. Black

We propose a new 3D model of the human body that is both realistic and part-based. The body is represented by a graphical model in which nodes of the graph correspond to body parts that can independently translate and ro…

Counterfactual Monotonic Knowledge Tracing for Assessing Students' Dynamic Mastery of Knowledge Concepts

2023-08-07 · Moyu Zhang, Xinning Zhu, Chunhong Zhang, Wenchen Qian 외

As the core of the Knowledge Tracking (KT) task, assessing students' dynamic mastery of knowledge concepts is crucial for both offline teaching and online educational applications. Since students' mastery of knowledge co…

counterfactualKnowledge Tracing