Beyond Trivial Counterfactual Explanations with Diverse Valuable Explanations
Explainability for machine learning models has gained considerable attention within the research community given the importance of deploying more reliable machine-learning systems. In computer vision applications, generative counterfactual methods indicate how to perturb a model's input to change its prediction, providing details about the model's decision-making. Current methods tend to generate trivial counterfactuals about a model's decisions, as they often suggest to exaggerate or remove the presence of the attribute being classified. For the machine learning practitioner, these types of counterfactuals offer little value, since they provide no new information about undesired model or data biases. In this work, we identify the problem of trivial counterfactual generation and we propose DiVE to alleviate it. DiVE learns a perturbation in a disentangled latent space that is constrained using a diversity-enforcing loss to uncover multiple valuable explanations about the model's prediction. Further, we introduce a mechanism to prevent the model from producing trivial explanations. Experiments on CelebA and Synbols demonstrate that our model improves the success rate of producing high-quality valuable explanations when compared to previous state-of-the-art methods. Code is available at https://github.com/ElementAI/beyond-trivial-explanations.
Code (4)
Tasks
AttributeBIG-bench Machine LearningcounterfactualDecision MakingDiversityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Beyond Trivial Counterfactual Generations with Diverse Valuable Explanations
Explainability of black-box predictive models has gained considerable attention within our research community given the importance of deploying more reliable machine-learning systems. Explanability can also be helpful f…
AttributecounterfactualDecision MakingTowards A Unified Information Bottleneck Framework for Time Series Explanations
Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into …
Beyond One-Size-Fits-All: Adapting Counterfactual Explanations to User Objectives
Explainable Artificial Intelligence (XAI) has emerged as a critical area of research aimed at enhancing the transparency and interpretability of AI systems. Counterfactual Explanations (CFEs) offer valuable insights into…
AllcounterfactualDecision MakingExplainable artificial intelligence+1Beyond Known Reality: Exploiting Counterfactual Explanations for Medical Research
The field of explainability in artificial intelligence (AI) has witnessed a growing number of studies and increasing scholarly interest. However, the lack of human-friendly and individual interpretations in explaining th…
counterfactualData AugmentationDecision MakingExplaining Machine Learning Classifiers through Diverse Counterfactual Explanations
Post-hoc explanations of machine learning models are crucial for people to understand and act on algorithmic predictions. An intriguing class of explanations is through counterfactuals, hypothetical examples that show pe…
BIG-bench Machine LearningcounterfactualDiversityPoint Processes