paper-with-me

홈 › Papers

Enforcing Interpretability and its Statistical Impacts: Trade-offs between Accuracy and Interpretability

2020-10-26 · Gintare Karolina Dziugaite, Shai Ben-David, Daniel M. Roy

To date, there has been no formal study of the statistical cost of interpretability in machine learning. As such, the discourse around potential trade-offs is often informal and misconceptions abound. In this work, we aim to initiate a formal study of these trade-offs. A seemingly insurmountable roadblock is the lack of any agreed upon definition of interpretability. Instead, we propose a shift in perspective. Rather than attempt to define interpretability, we propose to model the \emph{act} of \emph{enforcing} interpretability. As a starting point, we focus on the setting of empirical risk minimization for binary classification, and view interpretability as a constraint placed on learning. That is, we assume we are given a subset of hypothesis that are deemed to be interpretable, possibly depending on the data distribution and other aspects of the context. We then model the act of enforcing interpretability as that of performing empirical risk minimization over the set of interpretable hypotheses. This model allows us to reason about the statistical implications of enforcing interpretability, using known results in statistical learning theory. Focusing on accuracy, we perform a case analysis, explaining why one may or may not observe a trade-off between accuracy and interpretability when the restriction to interpretable classifiers does or does not come at the cost of some excess statistical risk. We close with some worked examples and some open problems, which we hope will spur further theoretical development around the tradeoffs involved in interpretability.

📄 PDF Abstract BibTeX arXiv:2010.13764

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationLearning TheoryMisconceptions

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

Coupling Agent-based Modeling and Life Cycle Assessment to Analyze Trade-offs in Resilient Energy Transitions

2025-11-10 · Beichen Zhang, Mohammed T. Zaki, Hanna Breunig, Newsha K. Ajami arxiv

Transitioning to sustainable and resilient energy systems requires navigating complex and interdependent trade-offs across environmental, social, and resource dimensions. Neglecting these trade-offs can lead to unintende…

Decision Making

Quantifying the Accuracy-Interpretability Trade-Off in Concept-Based Sidechannel Models

2025-10-07 · David Debot, Giuseppe Marra arxiv

Concept Bottleneck Models (CBNMs) are deep learning models that provide interpretability by enforcing a bottleneck layer where predictions are based exclusively on human-understandable concepts. However, this constraint …

Federated Unlearning: a Perspective of Stability and Fairness

2024-02-02 · Jiaqi Shao, Tao Lin, Xuanyu Cao, Bing Luo

This paper explores the multifaceted consequences of federated unlearning (FU) with data heterogeneity. We introduce key metrics for FU assessment, concentrating on verification, global stability, and local fairness, and…

Fairness

No-Free-Fairness: Fundamental Limits and Trade-offs in Learning Systems

2026-06-16 · Khoat Than arxiv

In this paper, we establish a set of theoretical impossibility results, termed the No-Free-Fairness theorems, that identify three fundamental sources of disparity in learning systems. First, we show that when a task exhi…

Counterfactual Generation with Knockoffs

2021-02-01 · Oana-Iuliana Popescu, Maha Shadaydeh, Joachim Denzler

Human interpretability of deep neural networks' decisions is crucial, especially in domains where these directly affect human lives. Counterfactual explanations of already trained neural networks can be generated by pert…

counterfactualVariable Selection