paper-with-me

홈 › Papers

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

2025-11-29 · Sitong Fang, Shiyi Hou, Kaile Wang, Boyuan Chen, Donghai Hong, Jiayi Zhou, Josef Dai, Yaodong Yang, Jiaming Ji arxiv

Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their performance leaps lie more insidious and destructive safety risks, namely deception. Unlike hallucination, which arises from insufficient capability and leads to mistakes, deception represents a deeper threat in which models deliberately mislead users through complex reasoning and insincere responses. As system capabilities advance, deceptive behaviours have spread from textual to multimodal settings, amplifying their potential harm. First and foremost, how can we monitor these covert multimodal deceptive behaviors? Nevertheless, current research remains almost entirely confined to text, leaving the deceptive risks of multimodal large language models unexplored. In this work, we systematically reveal and quantify multimodal deception risks, introducing MM-DeceptionBench, the first benchmark explicitly designed to evaluate multimodal deception. Covering six categories of deception, MM-DeceptionBench characterizes how models strategically manipulate and mislead through combined visual and textual modalities. On the other hand, multimodal deception evaluation is almost a blind spot in existing methods. Its stealth, compounded by visual-semantic ambiguity and the complexity of cross-modal reasoning, renders action monitoring and chain-of-thought monitoring largely ineffective. To tackle this challenge, we propose debate with images, a novel multi-agent debate monitor framework. By compelling models to ground their claims in visual evidence, this method substantially improves the detectability of deceptive strategies. Experiments show that it consistently increases agreement with human judgements across all tested models, boosting Cohen's kappa by 1.5x and accuracy by 1.25x on GPT-4o.

📄 PDF Abstract BibTeX arXiv:2512.00349

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Box of Lies: Multimodal Deception Detection in Dialogues

2019-06-01 · NAACL 2019 6 · Felix Soldner, Ver{\'o}nica P{\'e}rez-Rosas, Rada Mihalcea

Deception often takes place during everyday conversations, yet conversational dialogues remain largely unexplored by current work on automatic deception detection. In this paper, we address the task of detecting multimod…

Deception DetectionGeneral Classification

Introducing Representations of Facial Affect in Automated Multimodal Deception Detection

2020-08-31 · Leena Mathur, Maja J. Matarić

Automated deception detection systems can enhance health, justice, and security in society by helping humans detect deceivers in high-stakes situations across medical and legal domains, among others. This paper presents …

Deception DetectionEmotion Recognition

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

2025-07-12 · Santhosh Kumar Ravindran

Large language models (LLMs) aligned for safety through techniques like reinforcement learning from human feedback (RLHF) often exhibit emergent deceptive behaviors, where outputs appear compliant but subtly mislead or o…

Anomaly Detection

WOLF: Werewolf-based Observations for LLM Deception and Falsehoods

2025-12-09 · Mrinal Agarwal, Saad Rana, Theo Sundoro, Hermela Berhe 외 arxiv

Deception is a fundamental challenge for multi-agent reasoning: effective systems must strategically conceal information while detecting misleading behavior in others. Yet most evaluations reduce deception to static clas…

Affect-Aware Deep Belief Network Representations for Multimodal Unsupervised Deception Detection

2021-08-17 · Leena Mathur, Maja J Matarić

Automated systems that detect the social behavior of deception can enhance human well-being across medical, social work, and legal domains. Labeled datasets to train supervised deception detection models can rarely be co…

Deception Detection