paper-with-me

홈 › Papers

Framework for Evaluating Faithfulness of Local Explanations

2022-02-01 · Sanjoy Dasgupta, Nave Frost, Michal Moshkovitz

We study the faithfulness of an explanation system to the underlying prediction model. We show that this can be captured by two properties, consistency and sufficiency, and introduce quantitative measures of the extent to which these hold. Interestingly, these measures depend on the test-time data distribution. For a variety of existing explanation systems, such as anchors, we analytically study these quantities. We also provide estimators and sample complexity bounds for empirically determining the faithfulness of black-box explanation systems. Finally, we experimentally validate the new properties and estimators.

📄 PDF Abstract BibTeX arXiv:2202.00734

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Readability and Faithfulness of Concept-based Explanations

2024-04-29 · Meng Li, Haoran Jin, Ruixuan Huang, Zhihao Xu 외

With the growing popularity of general-purpose Large Language Models (LLMs), comes a need for more global explanations of model behaviors. Concept-based explanations arise as a promising avenue for explaining high-level …

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

2026-07-10 · Mohadeseh Mollapour, Koorosh Aslansefat, Zeinab Dehghani, Bhupesh Kumar Mishra 외 arxiv

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturba…

ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs

2026-03-19 · Abhinaba Basu, Pavan Chakraborty arxiv

Evaluating whether explanations faithfully reflect a model's reasoning remains an open problem. Existing benchmarks use single interventions without statistical testing, making it impossible to distinguish genuine faithf…

Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs

2024-09-18 · Christos Fragkathoulas, Odysseas S. Chlapanis

This paper introduces a novel task to assess the faithfulness of large language models (LLMs) using local perturbations and self-explanations. Many LLMs often require additional context to answer certain questions correc…

Natural Questions

A Causal Lens for Evaluating Faithfulness Metrics

2025-02-26 · Kerem Zaman, Shashank Srivastava

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's internal…

Decision MakingFact CheckingModel EditingObject Counting