paper-with-me

홈 › Papers

"How do I fool you?": Manipulating User Trust via Misleading Black Box Explanations

2019-11-15 · Himabindu Lakkaraju, Osbert Bastani

As machine learning black boxes are increasingly being deployed in critical domains such as healthcare and criminal justice, there has been a growing emphasis on developing techniques for explaining these black boxes in a human interpretable manner. It has recently become apparent that a high-fidelity explanation of a black box ML model may not accurately reflect the biases in the black box. As a consequence, explanations have the potential to mislead human users into trusting a problematic black box. In this work, we rigorously explore the notion of misleading explanations and how they influence user trust in black-box models. More specifically, we propose a novel theoretical framework for understanding and generating misleading explanations, and carry out a user study with domain experts to demonstrate how these explanations can be used to mislead users. Our work is the first to empirically establish how user trust in black box models can be manipulated via misleading explanations.

📄 PDF Abstract BibTeX arXiv:1911.06473

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fooling Partial Dependence via Data Poisoning

2021-05-26 · Hubert Baniecki, Wojciech Kretowicz, Przemyslaw Biecek

Many methods have been developed to understand complex predictive models and high expectations are placed on post-hoc model explainability. It turns out that such explanations are not robust nor trustworthy, and they can…

Data Poisoning

PETGEN: Personalized Text Generation Attack on Deep Sequence Embedding-based Classification Models

2021-09-14 · Bing He, Mustaque Ahamad, Srijan Kumar

What should a malicious user write next to fool a detection model? Identifying malicious users is critical to ensure the safety and integrity of internet platforms. Several deep learning-based detection models have been …

Adversarial AttackText Generation

CAAD 2018: Iterative Ensemble Adversarial Attack

2018-11-07 · Jiayang Liu, Weiming Zhang, Nenghai Yu

Deep Neural Networks (DNNs) have recently led to significant improvements in many fields. However, DNNs are vulnerable to adversarial examples which are samples with imperceptible perturbations while dramatically mislead…

Adversarial Attack

Unbiased Rectification for Sequential Recommender Systems Under Fake Orders

2026-01-24 · Qiyu Qin, Yichen Li, Haozhao Wang, Cheng Wang 외 arxiv

Fake orders pose increasing threats to sequential recommender systems by misleading recommendation results through artificially manipulated interactions, including click farming, context-irrelevant substitutions, and seq…

Computational EfficiencyData Augmentation

StyleFool: Fooling Video Classification Systems via Style Transfer

2022-03-30 · Yuxin Cao, Xi Xiao, Ruoxi Sun, Derui Wang 외

Video classification systems are vulnerable to adversarial attacks, which can create severe security problems in video verification. Current black-box attacks need a large number of queries to succeed, resulting in high …

Adversarial AttackClassificationDenoisingStyle Transfer+1