paper-with-me

홈 › Papers

Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?

2025-10-28 · Teague McMillan, Gabriele Dominici, Martin Gjoreski, Marc Langheinrich arxiv

Large Language Models (LLMs) often produce explanations that do not faithfully reflect the factors driving their predictions. In healthcare settings, such unfaithfulness is especially problematic: explanations that omit salient clinical cues or mask spurious shortcuts can undermine clinician trust and lead to unsafe decision support. We study how inference and training-time choices shape explanation faithfulness, focusing on factors practitioners can control at deployment. We evaluate three LLMs (GPT-4.1-mini, LLaMA 70B, LLaMA 8B) on two datasets-BBQ (social bias) and MedQA (medical licensing questions), and manipulate the number and type of few-shot examples, prompting strategies, and training procedure. Our results show: (i) both the quantity and quality of few-shot examples significantly impact model faithfulness; (ii) faithfulness is sensitive to prompting design; (iii) the instruction-tuning phase improves measured faithfulness on MedQA. These findings offer insights into strategies for enhancing the interpretability and trustworthiness of LLMs in sensitive domains.

📄 PDF Abstract BibTeX arXiv:2510.24236

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

2025-10-31 · Austin Meek, Eitan Sprejer, Iván Arcuschin, Austin J. Brockmeier 외 arxiv

Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT is a direct window into what the model is…

RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models

2026-02-19 · Yunseok Han, Yejoon Lee, Jaeyoung Do arxiv

Large Reasoning Models (LRMs) exhibit strong performance, yet often produce rationales that sound plausible but fail to reflect their true decision process, undermining reliability and trust. We introduce a formal framew…

GRACE: Step-Level Benchmark for Faithful Reasoning over Context

2026-06-15 · Hoang Pham, Dong Le, Anh Tuan Luu arxiv

Many reasoning tasks require models to reason over input context, from document-grounded question answering to rule-based deduction. Chain-of-Thought (CoT) prompting produces traces that appear transparent, yet individua…

Reinforcement LearningQuestion Answering

Zero-Shot Textual Explanations via Translating Decision-Critical Features

2025-12-08 · Toshinori Yamauchi, Hiroshi Kera, Kazuhiko Kawamoto arxiv

Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language models can generate captions but are designed for general visual understa…

Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models

2024-02-07 · Chirag Agarwal, Sree Harsha Tanneru, Himabindu Lakkaraju

Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermed…

Decision Making