paper-with-me

홈 › Papers

On the Sensitivity and Stability of Model Interpretations

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions. Despite the surge of new interpretation methods, it remains an open problem how to define and quantitatively measure the faithfulness of interpretations, i.e., to what extent interpretations reflect the reasoning process by a model. We propose two new criteria, sensitivity and stability, that provide complementary notions of faithfulness to the existed removal-based criteria. Our results show that the conclusion for how faithful interpretations are could vary substantially based on different notions. Motivated by the desiderata of sensitivity and stability, we introduce a new class of interpretation methods that adopt techniques from adversarial robustness. Empirical results show that our proposed methods are effective under the new criteria and overcome limitations of gradient-based methods on removal-based criteria. Besides text classification, we also apply interpretation methods and metrics to dependency parsing. Our results shed light on understanding the diverse set of interpretations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessDependency ParsingmodelSensitivitytext-classificationText Classification

Similar Papers 제목 키워드 기반

On the Sensitivity and Stability of Model Interpretations in NLP

2021-04-18 · ACL 2022 5 · Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, Kai-Wei Chang

Recent years have witnessed the emergence of a variety of post-hoc interpretations that aim to uncover how natural language processing (NLP) models make predictions. Despite the surge of new interpretation methods, it re…

Adversarial RobustnessDependency ParsingSensitivitytext-classification+1

Are machine learning interpretations reliable? A stability study on global interpretations

2025-05-21 · Luqin Gan, Tarek M. Zikry, Genevera I. Allen

As machine learning systems are increasingly used in high-stakes domains, there is a growing emphasis placed on making them interpretable to improve trust in these systems. In response, a range of interpretable machine l…

Interpretable Machine Learning

Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing

2025-10-05 · Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong arxiv

Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing. We present a diagnostic analy…

Disentangling Ambiguity from Instability in Large Language Models: A Clinical Text-to-SQL Case Study

2026-02-12 · Angelo Ziletti, Leonardo D'Ambrosi arxiv

Deploying large language models for clinical Text-to-SQL requires distinguishing two qualitatively different causes of output diversity: (i) input ambiguity that should trigger clarification, and (ii) model instability t…

Interpretable Machine Learning for Football Performance Analysis: Evidence of Limited Transferability from Elite Leagues to University Competition

2026-05-11 · Yu-Fang Tsai, Yu-Jen Chen, Kok-Hua Tan, Sheng-Chieh Huang 외 arxiv

Machine learning has become increasingly prevalent in football performance analysis, yet most studies prioritize predictive accuracy while implicitly assuming that learned performance determinants and their interpretatio…

Interpretable Machine Learning