paper-with-me

홈 › Papers

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

2026-05-19 · Somnath Banerjee, Pranav Jha, Rima Hazra, Animesh Mukherjee arxiv

LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans as evidence alongside a generated rationale'' and uncover a systematic trade-off: English-pivot explanations can achieve higher span agreement with human rationales while their evidence becomes less causally grounded in the model's prediction, as measured by both comprehensiveness and sufficiency. Across 3 tasks, 5~languages, and 2~multilingual LLM families, we find that English explanations frequently produce fluent but loosely anchored rationales, with comprehensiveness degrading by up to 5.7x relative to native-language conditions - even as task accuracy remains stable across settings. For socially nuanced classification, English pivots also fail to preserve pragmatic cues, reducing both faithfulness and span agreement. We recommend auditing explanations in the input language, reporting multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than faithful decision traces.

📄 PDF Abstract BibTeX arXiv:2605.19274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparing interpretation methods in mental state decoding analyses with deep learning models

2022-05-31 · Armin W. Thomas, Christopher Ré, Russell A. Poldrack

Deep learning (DL) models find increasing application in mental state decoding, where researchers seek to understand the mapping between mental states (e.g., perceiving fear or joy) and brain activity by identifying thos…

Explainable artificial intelligence

Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models

2024-02-07 · Chirag Agarwal, Sree Harsha Tanneru, Himabindu Lakkaraju

Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermed…

Decision Making

Exploring the Trade-off Between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales

2024-04-03 · Lucas E. Resck, Marcos M. Raimundo, Jorge Poco

Saliency post-hoc explainability methods are important tools for understanding increasingly complex NLP models. While these methods can reflect the model's reasoning, they may not align with human intuition, making the e…

Contrastive LearningHate Speech DetectionSentiment Classificationtext-classification+1

A Multilingual Perspective Towards the Evaluation of Attribution Methods

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Most evaluations of attribution methods focus on the English language. In this work, we present a multilingual approach for evaluating attribution methods for the Natural Language Inference (NLI) task in terms of plausib…

Natural Language Inference

A Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference

2022-04-11 · Kerem Zaman, Yonatan Belinkov

Most evaluations of attribution methods focus on the English language. In this work, we present a multilingual approach for evaluating attribution methods for the Natural Language Inference (NLI) task in terms of faithfu…

Natural Language Inference