paper-with-me

홈 › Papers

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

2026-04-09 · Wenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu, Xuehe Wang, Edith Cheuk Han Ngai arxiv

In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However, coherent reasoning can still violate logical or evidential constraints, allowing unsupported beliefs repeatedly stored and propagated across decision steps, leading to systematic behavioral drift in long-horizon agentic systems. Most existing strategies rely on the consensus mechanism, conflating agreement with faithfulness. In this paper, inspired by the vulnerability of unfaithful intermediate reasoning trajectories, we propose \textbf{S}elf-\textbf{A}udited \textbf{Ve}rified \textbf{R}easoning (\textsc{SAVeR}), a novel framework that enforces verification over internal belief states within the agent before action commitment, achieving faithful reasoning. Concretely, we structurally generate persona-based diverse candidate beliefs for selection under a faithfulness-relevant structure space. To achieve reasoning faithfulness, we perform adversarial auditing to localize violations and repair through constraint-guided minimal interventions under verifiable acceptance criteria. Extensive experiments on six benchmark datasets demonstrate that our approach consistently improves reasoning faithfulness while preserving competitive end-task performance.

📄 PDF Abstract BibTeX arXiv:2604.08401

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit

2026-01-09 · Junda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu 외 arxiv

LLM agents operating in open environments face escalating risks from indirect prompt injection, particularly within the tool stream where manipulated metadata and runtime feedback hijack execution flow. Existing defenses…

Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking

2025-12-13 · Rheeya Uppaal, Phu Mon Htut, Min Bai, Nikolaos Pappas 외 arxiv

Reasoning-augmented vision language models (VLMs) generate explicit chains of thought that promise greater capability and transparency but also introduce new failure modes: models may reach correct answers via visually u…

Multimodal Reasoning

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

2026-08-29 · Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan 외 hf

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents…

Monitor-Generate-Verify (MGV): Formalising Metacognitive Theory for Language Model Reasoning

2025-11-06 · Nick Oh, Fernand Gobet arxiv

Test-time reasoning architectures such as those following the Generate-Verify paradigm, where a model iteratively refines or verifies its own generated outputs, prioritise generation and verification but exclude the moni…

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

2026-07-09 · Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls 외 arxiv

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal repres…