paper-with-me

홈 › Papers

FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness

2025-09-10 · Anand Swaroop, Akshat Nallani, Saksham Uboweja, Adiliia Uzdenova, Michael Nguyen, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Vasu Sharma, Maheep Chaudhary arxiv

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for improving large language model performance on complex tasks, but recent work shows that reasoning steps often fail to causally influence the final answer, creating brittle and untrustworthy outputs. Prior approaches focus primarily on measuring faithfulness, while methods for systematically improving it remain limited. We introduce Faithful Reasoning via Intervention Training (FRIT), a scalable alignment method that trains models to produce causally consistent reasoning by learning from systematically corrupted examples. FRIT generates synthetic training data by intervening on individual reasoning steps in model-generated CoTs, creating faithful/unfaithful pairs that highlight when reasoning breaks down. We then apply Direct Preference Optimization to teach models to prefer causally consistent reasoning paths. Evaluating on Qwen3-8B and Mistral-7B-v0.1 across factual and symbolic reasoning tasks, FRIT increases faithful reasoning by $3.4$ percentage points for Mistral on GSM8K while improving accuracy by $7.6$ percentage points. Our approach provides the first scalable, supervision-free method for training language models to produce more reliable and interpretable reasoning, addressing a critical gap between reasoning performance and trustworthiness. We release our code at \href{https://github.com/Anut-py/frit}.

📄 PDF Abstract BibTeX arXiv:2509.13334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Local Causal Attribution of Chain-of-Thought Reasoning

2026-06-20 · Dennis Wei, Yannis Belkhiter, Erik Miehling, Radu Marinescu arxiv

Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In this work, we take a local approach toward this goal by analyzing the c…

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

2026-04-23 · Qinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin 외 arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLV…

Reinforcement Learning

Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment

2024-03-05 · Congzhi Zhang, Linhai Zhang, Jialong Wu, Yulan He 외

Despite the notable advancements of existing prompting methods, such as In-Context Learning and Chain-of-Thought for Large Language Models (LLMs), they still face challenges related to various biases. Traditional debiasi…

Contrastive LearningData AugmentationIn-Context LearningLanguage Modeling+2

CHECKWHY: Causal Fact Verification via Argument Structure

2024-08-20 · Jiasheng Si, Yibo Zhao, Yingjie Zhu, Haiyang Zhu 외

With the growing complexity of fact verification tasks, the concern with "thoughtful" reasoning capabilities is increasing. However, recent fact verification benchmarks mainly focus on checking a narrow scope of semantic…

Fact VerificationLogical Reasoning

Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

2026-06-11 · Daniel Scalena, Sara Candussio, Luca Bortolussi, Elisabetta Fersini 외 arxiv

Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the causal influence of individual steps on the final answer poorly understood. We estimate each step's causal …