paper-with-me

홈 › Papers

DualCF: Efficient Model Extraction Attack from Counterfactual Explanations

2022-05-13 · Yongjie Wang, Hangwei Qian, Chunyan Miao

Cloud service providers have launched Machine-Learning-as-a-Service (MLaaS) platforms to allow users to access large-scale cloudbased models via APIs. In addition to prediction outputs, these APIs can also provide other information in a more human-understandable way, such as counterfactual explanations (CF). However, such extra information inevitably causes the cloud models to be more vulnerable to extraction attacks which aim to steal the internal functionality of models in the cloud. Due to the black-box nature of cloud models, however, a vast number of queries are inevitably required by existing attack strategies before the substitute model achieves high fidelity. In this paper, we propose a novel simple yet efficient querying strategy to greatly enhance the querying efficiency to steal a classification model. This is motivated by our observation that current querying strategies suffer from decision boundary shift issue induced by taking far-distant queries and close-to-boundary CFs into substitute model training. We then propose DualCF strategy to circumvent the above issues, which is achieved by taking not only CF but also counterfactual explanation of CF (CCF) as pairs of training samples for the substitute model. Extensive and comprehensive experimental evaluations are conducted on both synthetic and real-world datasets. The experimental results favorably illustrate that DualCF can produce a high-fidelity model with fewer queries efficiently and effectively.

📄 PDF Abstract BibTeX arXiv:2205.06504

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualCounterfactual ExplanationmodelModel extraction

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Model extraction from counterfactual explanations

2020-09-03 · Ulrich Aïvodji, Alexandre Bolot, Sébastien Gambs

Post-hoc explanation techniques refer to a posteriori methods that can be used to explain how black-box machine learning models produce their outcomes. Among post-hoc explanation techniques, counterfactual explanations a…

counterfactualmodelModel extraction

Watermarking Counterfactual Explanations

2024-05-29 · Hangzhi Guo, Firdaus Ahmed Choudhury, Tinghua Chen, Amulya Yadav

Counterfactual (CF) explanations for ML model predictions provide actionable recourse recommendations to individuals adversely impacted by predicted outcomes. However, despite being preferred by end-users, CF explanation…

counterfactualExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Model extraction

The privacy issue of counterfactual explanations: explanation linkage attacks

2022-10-21 · Sofie Goethals, Kenneth Sörensen, David Martens

Black-box machine learning models are being used in more and more high-stakes domains, which creates a growing need for Explainable AI (XAI). Unfortunately, the use of XAI in machine learning introduces new privacy risks…

counterfactualExplainable Artificial Intelligence (XAI)

Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations

2024-04-04 · Fatima Ezzeddine, Omran Ayoub, Silvia Giordano

In recent years, there has been a notable increase in the deployment of machine learning (ML) models as services (MLaaS) across diverse production software applications. In parallel, explainable AI (XAI) continues to evo…

counterfactualKnowledge DistillationModel extraction

Tabular Diffusion based Actionable Counterfactual Explanations for Network Intrusion Detection

2025-07-23 · Vinura Galwaduge, Jagath Samarabandu arxiv

Modern network intrusion detection systems (NIDS) frequently utilize the predictive power of complex deep learning models. However, the "black-box" nature of such deep learning methods adds a layer of opaqueness that hin…

Network Intrusion Detection