paper-with-me

Papers

Manipulation Risks in Explainable AI: The Implications of the Disagreement Problem

2023-06-24 · Sofie Goethals, David Martens, Theodoros Evgeniou

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of Explainable AI (XAI) has emerged in response. However, it faces a significant challenge known as the disagreement problem, where multiple explanations are possible for the same AI decision or prediction. While the existence of the disagreement problem is acknowledged, the potential implications associated with this problem have not yet been widely studied. First, we provide an overview of the different strategies explanation providers could deploy to adapt the returned explanation to their benefit. We make a distinction between strategies that attack the machine learning model or underlying data to influence the explanations, and strategies that leverage the explanation phase directly. Next, we analyse several objectives and concrete scenarios the providers could have to engage in this behavior, and the potential dangerous consequences this manipulative behavior could have on society. We emphasize that it is crucial to investigate this issue now, before these methods are widely implemented, and propose some mitigation strategies.

📄 PDF Abstract BibTeX arXiv:2306.13885

Code (0)

등록된 구현이 없습니다.

Tasks

Explainable Artificial Intelligence (XAI)

Similar Papers 제목 키워드 기반

EXAGREE: Towards Explanation Agreement in Explainable Machine Learning

2024-11-04 · Sichao Li, Quanling Deng, Amanda S. Barnard

Explanations in machine learning are critical for trust, transparency, and fairness. Yet, complex disagreements among these explanations limit the reliability and applicability of machine learning models, especially in h…

Fairness

Two-Person Bargaining when the Disagreement Point is Private Information

2022-11-13 · Eric van Damme, Xu Lang

We consider two-person bargaining problems in which (only) the disagreement outcome is private (and possibly correlated) information and it is common knowledge that disagreement is inefficient. We show that if the Pareto…

Explainable News Summarization -- Analysis and mitigation of Disagreement Problem

2024-10-24 · Seema Aswani, Sujala D. Shetty

Explainable AI (XAI) techniques for text summarization provide valuable understanding of how the summaries are generated. Recent studies have highlighted a major challenge in this area, known as the disagreement problem.…

Extreme SummarizationNews SummarizationText Summarization

The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective

2022-02-03 · Satyapriya Krishna, Tessa Han, Alex Gu, Steven Wu 외

As various post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to develop a deeper understanding of whether and when the explanations outpu…

BIG-bench Machine Learning

ExTax: Explainable Disinformation Detection via Persuasion, Emotion, and Narrative Role Taxonomies

2026-05-26 · Shang Luo, Yingguang Yang, Zhenchen Sun, Yang Liu 외 arxiv

The democratization of LLMs has accelerated the generation and circulation of highly fluent disinformation, making traditional syntax-semantic verification increasingly insufficient. Such deception rarely relies solely o…