paper-with-me

홈 › Papers

A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations

2026-04-19 · Marcin Kostrzewa, Maciej Zięba, Jerzy Stefanowski arxiv

Counterfactual explanations (CFEs) are essential for interpreting black-box models, yet they often become invalid when models are slightly changed. Existing methods for generating robust CFEs are often limited to specific types of models, require costly tuning, or inflexible robustness controls. We propose a novel approach that jointly models the data distribution and the space of plausible model decisions to ensure robustness to model changes. Using a probabilistic consensus over a model ensemble, we train a conditional normalizing flow that captures the data density under varying levels of classifier agreement. At inference time, a single interpretable parameter controls the robustness level; it specifies the minimum fraction of models that should agree on the target class without retraining the generative model. Our method effectively pushes CFEs toward regions that are both plausible and stable across model changes. Experimental results demonstrate that our approach achieves superior empirical robustness while also maintaining good performance across other evaluation measures.

📄 PDF Abstract BibTeX arXiv:2604.17494

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Why Did This Model Forecast This Future? Information-Theoretic Saliency for Counterfactual Explanations of Probabilistic Regression Models

2023-09-21 · NeurIPS 2023 11

We propose a post hoc saliency-based explanation framework for counterfactual reasoning in probabilistic multivariate time-series forecasting (regression) settings. Building upon Miller's framework of explanations derive…

Provably Robust Bayesian Counterfactual Explanations under Model Changes

2026-01-23 · Jamie Duell, Xiuyi Fan arxiv

Counterfactual explanations (CEs) offer interpretable insights into machine learning predictions by answering ``what if?" questions. However, in real-world settings where models are frequently updated, existing counterfa…

DifCluE: Generating Counterfactual Explanations with Diffusion Autoencoders and modal clustering

2025-02-17 · Suparshva Jain, Amit Sangroya, Lovekesh Vig

Generating multiple counterfactual explanations for different modes within a class presents a significant challenge, as these modes are distinct yet converge under the same classification. Diffusion probabilistic models …

Clusteringcounterfactual

Gradient-based Counterfactual Explanations using Tractable Probabilistic Models

2022-05-16 · Xiaoting Shao, Kristian Kersting

Counterfactual examples are an appealing class of post-hoc explanations for machine learning models. Given input $x$ of class $y_1$, its counterfactual is a contrastive example $x^\prime$ of another class $y_0$. Current …

counterfactual

Disagreement amongst counterfactual explanations: How transparency can be deceptive

2023-04-25 · Dieter Brughmans, Lissa Melis, David Martens

Counterfactual explanations are increasingly used as an Explainable Artificial Intelligence (XAI) technique to provide stakeholders of complex machine learning algorithms with explanations for data-driven decisions. The …

counterfactualDecision MakingDiversityExplainable artificial intelligence+1