paper-with-me

Papers

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

2025-10-26 · Haoyi Qiu, Yilun Zhou, Pranav Narayanan Venkit, Kung-Hsiang Huang, Jiaxin Zhang, Nanyun Peng, Chien-Sheng Wu arxiv

As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored the dynamics of Agent-to-Agent (A2A) persuasion, the rise of Vision-Language Models (VLMs) introduces a more complex challenge: multimodal content conveys richer information while integrating subtle, hard-to-detect persuasive cues. To study this vulnerability, we present MMPersuade, a unified framework and dataset for A2A multimodal persuasion. We model interactions between a persuader agent, which leverages images and psychological strategies, and a persuadee VLM. Our benchmark spans commercial, subjective and behavioral, and adversarial contexts, and evaluates persuasion via function-calling that capture behavioral shifts beyond verbal responses. Experiments on six VLMs reveal three findings: (1) multimodal inputs consistently outperform text-only persuasion, with raw visual signals uniquely increasing susceptibility in adversarial settings by bypassing text-activated safety defenses; (2) persuadee vulnerability is highly domain- and format-dependent, with realistic and community-style formats driving susceptibility in commercial settings while different formats dominate in adversarial ones; and (3) psychological strategy efficacy varies with context and model architecture, as more capable models resist benign persuasion yet become more susceptible under adversarial multimodal inputs. Our framework provides a foundation for building more robust and aligned VLMs in multi-agent environments.

📄 PDF Abstract BibTeX arXiv:2510.22768

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations

2025-09-30 · Keyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren 외 arxiv

When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM predictions with natural language explanations can signal which model p…

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

2025-10-20 · Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo 외 arxiv

Vision-Language Models (VLMs) achieve strong results on multimodal tasks such as visual question answering, yet they can still fail even when the correct visual evidence is present. In this work, we systematically invest…

Visual Question Answering

Believing is Seeing: Unobserved Object Detection using Generative Models

2024-10-08 · CVPR 2025 1 · Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome

Can objects that are not visible in an image -- but are in the vicinity of the camera -- be detected? This study introduces the novel tasks of 2D, 2.5D and 3D unobserved object detection for predicting the location of ne…

Objectobject-detectionObject Detection

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

2025-08-22 · Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello 외 arxiv

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion…

Emotion Recognition

Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

2025-11-15 · Pinxue Guo, Chongruo Wu, Xinyu Zhou, Lingyi Hong 외 arxiv

Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensur…

Visual Grounding