A Learning Theoretic Perspective on Local Explainability
In this paper, we explore connections between interpretable machine learning and learning theory through the lens of local approximation explanations. First, we tackle the traditional problem of performance generalization and bound the test-time accuracy of a model using a notion of how locally explainable it is. Second, we explore the novel problem of explanation generalization which is an important concern for a growing class of finite sample-based local approximation explanations. Finally, we validate our theoretical results empirically and show that they reflect what can be seen in practice.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningInterpretable Machine LearningLearning TheorySimilar Papers 제목 키워드 기반
AI Explainability for Power Electronics: From a Lipschitz Continuity Perspective
Lifecycle management of power converters continues to thrive with emerging artificial intelligence (AI) solutions, yet AI mathematical explainability remains unexplored in power electronics (PE) community. The lack of th…
Fault DiagnosisThe Limits of AI Explainability: An Algorithmic Information Theory Approach
This paper establishes a theoretical foundation for understanding the fundamental limits of AI explainability through algorithmic information theory. We formalize explainability as the approximation of complex models by …
Automated Theorem ProvingThe geometry of BERT
Transformer neural networks, particularly Bidirectional Encoder Representations from Transformers (BERT), have shown remarkable performance across various tasks such as classification, text summarization, and question an…
Question AnsweringText SummarizationProbabilistic Lipschitzness and the Stable Rank for Comparing Explanation Models
Explainability models are now prevalent within machine learning to address the black-box nature of neural networks. The question now is which explainability model is most effective. Probabilistic Lipschitzness has demons…
Unifying Formal Explanations: A Complexity-Theoretic Perspective
Previous work has explored the computational complexity of deriving two fundamental types of explanations for ML model predictions: (1) *sufficient reasons*, which are subsets of input features that, when fixed, determin…