paper-with-me

홈 › Papers

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making

2026-02-03 · Shutong Fan, Lan Zhang, Xiaoyong Yuan arxiv

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and act on model recommendations. Large Language Models (LLMs) generate fluent natural-language explanations that shape how users perceive and trust AI outputs, revealing a new attack surface at the cognitive layer: the communication channel between AI and its users. We introduce adversarial explanation attacks (AEAs), where an attacker manipulates the framing of LLM-generated explanations to modulate human trust in incorrect outputs. We formalize this behavioral threat through the trust miscalibration gap, a metric that captures the difference in human trust between benign and adversarial explanations. Using this metric as a lens, we highlight a behavioral risk where persuasive explanation framing can preserve user trust even when the underlying AI prediction is wrong. To characterize this threat, we conducted a human study with over 200 participants, systematically varying four dimensions of explanation framing: reasoning mode, evidence type, communication style, and presentation format. Our findings show that users report nearly identical trust for adversarial and benign explanations, with adversarial explanations preserving the vast majority of benign trust despite being incorrect. The most vulnerable cases arise when AEAs closely resemble expert communication, combining authoritative evidence, neutral tone, and domain-appropriate reasoning. Vulnerability is highest on hard tasks, in fact-driven domains, and among participants who are less formally educated, younger, or highly trusting of AI.

📄 PDF Abstract BibTeX arXiv:2602.04003

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Overcoming Adversarial Attacks for Human-in-the-Loop Applications

2023-06-09 · Ryan McCoppin, Marla Kennedy, Platon Lukyanenko, Sean Kennedy

Including human analysis has the potential to positively affect the robustness of Deep Neural Networks and is relatively unexplored in the Adversarial Machine Learning literature. Neural network visual explanation maps h…

Uncertainty-Aware SAR ATR: Defending Against Adversarial Attacks via Bayesian Neural Networks

2024-03-27 · Tian Ye, Rajgopal Kannan, Viktor Prasanna, Carl Busart

Adversarial attacks have demonstrated the vulnerability of Machine Learning (ML) image classifiers in Synthetic Aperture Radar (SAR) Automatic Target Recognition (ATR) systems. An adversarial attack can deceive the class…

Adversarial AttackDecision Makingimage-classificationImage Classification

Resilience of Bayesian Layer-Wise Explanations under Adversarial Attacks

2021-02-22 · Ginevra Carbone, Guido Sanguinetti, Luca Bortolussi

We consider the problem of the stability of saliency-based explanations of Neural Network predictions under adversarial attacks in a classification task. Saliency interpretations of deterministic Neural Networks are rema…

General Classification

Impact of Adversarial Attacks on Deep Learning Model Explainability

2024-12-15 · Gazi Nazia Nur, Mohammad Ahnaf Sadat

In this paper, we investigate the impact of adversarial attacks on the explainability of deep learning models, which are commonly criticized for their black-box nature despite their capacity for autonomous feature extrac…

Decision MakingDeep Learning

Adversarial Counterfactual Visual Explanations

2023-03-17 · CVPR 2023 1 · Guillaume Jeanneret, Loïc Simon, Frédéric Jurie

Counterfactual explanations and adversarial attacks have a related goal: flipping output labels with minimal perturbations regardless of their characteristics. Yet, adversarial attacks cannot be used directly in a counte…

counterfactualCounterfactual ExplanationDenoising