paper-with-me

홈 › Papers

Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution

2025-12-11 · Jonathan Kamp, Roos Bakker, Dominique Blok arxiv

Good quality explanations strengthen the understanding of language models and data. Feature attribution methods, such as Integrated Gradient, are a type of post-hoc explainer that can provide token-level insights. However, explanations on the same input may vary greatly due to underlying biases of different methods. Users may be aware of this issue and mistrust their utility, while unaware users may trust them inadequately. In this work, we delve beyond the superficial inconsistencies between attribution methods, structuring their biases through a model- and method-agnostic framework of three evaluation metrics. We systematically assess both lexical and position bias (what and where in the input) for two transformers; first, in a controlled, pseudo-random classification task on artificial data; then, in a semi-controlled causal relation detection task on natural data. We find a trade-off between lexical and position biases in our model comparison, with models that score high on one type score low on the other. We also find signs that anomalous explanations are more likely to be biased.

📄 PDF Abstract BibTeX arXiv:2512.11108

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Open the Black Box Data-Driven Explanation of Black Box Decision Systems

2018-06-26 · Dino Pedreschi, Fosca Giannotti, Riccardo Guidotti, Anna Monreale 외

Black box systems for automated decision making, often based on machine learning over (big) data, map a user's features into a class or a score without exposing the reasons why. This is problematic not only for lack of t…

Decision Making

Revealing Hidden Context Bias in Segmentation and Object Detection through Concept-specific Explanations

2022-11-21 · Maximilian Dreyer, Reduan Achtibat, Thomas Wiegand, Wojciech Samek 외

Applying traditional post-hoc attribution methods to segmentation or object detection predictors offers only limited insights, as the obtained feature attribution maps at input level typically resemble the models' predic…

Explainable artificial intelligenceobject-detectionObject DetectionSegmentation

Beneath the Surface: How Large Language Models Reflect Hidden Bias

2025-02-27 · Jinhao Pan, Chahat Raj, Ziyu Yao, Ziwei Zhu

The exceptional performance of Large Language Models (LLMs) often comes with the unintended propagation of social biases embedded in their training data. While existing benchmarks evaluate overt bias through direct term …

On Explaining Proxy Discrimination and Unfairness in Individual Decisions Made by AI Systems

2025-09-30 · Belona Sonna, Alban Grastien arxiv

Artificial intelligence (AI) systems in high-stakes domains raise concerns about proxy discrimination, unfairness, and explainability. Existing audits often fail to reveal why unfairness arises, particularly when rooted …

Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models

2026-01-30 · Xilin Gong, Shu Yang, Zehua Cao, Lynne Billard 외 arxiv

Large Language Models (LLMs) have demonstrated strong capabilities for hidden representation interpretation through Patchscopes, a framework that uses LLMs themselves to generate human-readable explanations by decoding f…