paper-with-me

홈 › Papers

SHLIME: Foiling adversarial attacks fooling SHAP and LIME

2025-08-14 · Sam Chauhan, Estelle Duguet, Karthik Ramakrishnan, Hugh Van Deventer, Jack Kruger, Ranjan Subbaraman arxiv

Post hoc explanation methods, such as LIME and SHAP, provide interpretable insights into black-box classifiers and are increasingly used to assess model biases and generalizability. However, these methods are vulnerable to adversarial manipulation, potentially concealing harmful biases. Building on the work of Slack et al. (2020), we investigate the susceptibility of LIME and SHAP to biased models and evaluate strategies for improving robustness. We first replicate the original COMPAS experiment to validate prior findings and establish a baseline. We then introduce a modular testing framework enabling systematic evaluation of augmented and ensemble explanation approaches across classifiers of varying performance. Using this framework, we assess multiple LIME/SHAP ensemble configurations on out-of-distribution models, comparing their resistance to bias concealment against the original methods. Our results identify configurations that substantially improve bias detection, highlighting their potential for enhancing transparency in the deployment of high-stakes machine learning systems.

📄 PDF Abstract BibTeX arXiv:2508.11053

Code (0)

등록된 구현이 없습니다.

Tasks

Bias Detection

Similar Papers 제목 키워드 기반

Unified Adversarial Patch for Cross-modal Attacks in the Physical World

2023-07-15 · ICCV 2023 1 · Xingxing Wei, Yao Huang, Yitong Sun, Jie Yu

Recently, physical adversarial attacks have been presented to evade DNNs-based object detectors. To ensure the security, many scenarios are simultaneously deployed with visible sensors and infrared sensors, leading to th…

Fooling SHAP with Output Shuffling Attacks

2024-08-12 · Jun Yuan, Aritra Dasgupta

Explainable AI~(XAI) methods such as SHAP can help discover feature attributions in black-box models. If the method reveals a significant attribution from a ``protected feature'' (e.g., gender, race) on the model output,…

Adversarial Fooling Beyond "Flipping the Label"

2020-04-27 · Konda Reddy Mopuri, Vaisakh Shaj, R. Venkatesh Babu

Recent advancements in CNNs have shown remarkable achievements in various CV/AI applications. Though CNNs show near human or better than human performance in many critical tasks, they are quite vulnerable to adversarial …

Robust Android Malware Detection System against Adversarial Attacks using Q-Learning

2021-01-27 · Hemant Rathore, Sanjay K. Sahay, Piyush Nikam, Mohit Sewak

The current state-of-the-art Android malware detection systems are based on machine learning and deep learning models. Despite having superior performance, these models are susceptible to adversarial attacks. Therefore i…

Adversarial DefenseAndroid Malware DetectionBIG-bench Machine LearningMalware Detection+4

Unified Adversarial Patch for Visible-Infrared Cross-modal Attacks in the Physical World

2023-07-27 · Xingxing Wei, Yao Huang, Yitong Sun, Jie Yu

Physical adversarial attacks have put a severe threat to DNN-based object detectors. To enhance security, a combination of visible and infrared sensors is deployed in various scenarios, which has proven effective in disa…