paper-with-me

홈 › Papers

From Sufficiency to Reflection: Reinforcement-Guided Thinking Quality in Retrieval-Augmented Reasoning for LLMs

2025-07-30 · Jie He, Victor Gutiérrez-Basulto, Jeff Z. Pan arxiv

Reinforcement learning-based retrieval-augmented generation (RAG) methods enhance the reasoning abilities of large language models (LLMs). However, most rely only on final-answer rewards, overlooking intermediate reasoning quality. This paper analyzes existing RAG reasoning models and identifies three main failure patterns: (1) information insufficiency, meaning the model fails to retrieve adequate support; (2) faulty reasoning, where logical or content-level flaws appear despite sufficient information; and (3) answer-reasoning inconsistency, where a valid reasoning chain leads to a mismatched final answer. We propose TIRESRAG-R1, a novel framework using a think-retrieve-reflect process and a multi-dimensional reward system to improve reasoning and stability. TIRESRAG-R1 introduces: (1) a sufficiency reward to encourage thorough retrieval; (2) a reasoning quality reward to assess the rationality and accuracy of the reasoning chain; and (3) a reflection reward to detect and revise errors. It also employs a difficulty-aware reweighting strategy and training sample filtering to boost performance on complex tasks. Experiments on four multi-hop QA datasets show that TIRESRAG-R1 outperforms prior RAG methods and generalizes well to single-hop tasks. The code and data are available at: https://github.com/probe2/TIRESRAG-R1.

📄 PDF Abstract BibTeX arXiv:2507.22716

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning

2026-04-08 · Yang Xiang, Yixin Ji, Ruotao Xu, Dan Qiao 외 arxiv

Large reasoning models (LRMs) have achieved remarkable performance in complex reasoning tasks, driven by their powerful inference-time scaling capability. However, LRMs often suffer from overthinking, which results in su…

SuCo: Sufficiency-guided Continuous Adaptive Reasoning

2026-06-16 · Jiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang 외 arxiv

Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simple queries. Existing efforts to mitigate …

Reinforcement Learning

Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression

2025-08-07 · Jiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen 외 arxiv

Recent Large Reasoning Language Models (LRLMs) employ long chain-of-thought reasoning with complex reflection behaviors, typically signaled by specific trigger words (e.g., "Wait" and "Alternatively") to enhance performa…

VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

2025-04-10 · Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu 외

Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such…

MathMultimodal Reasoning

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model

2025-06-23 · Xu Wan, Wei Wang, Wenyue Xu, Wotao Yin 외

Reinforcement Learning (RL)-based post-training has significantly advanced the complex reasoning capabilities of language models, fostering sophisticated self-reflection processes. However, this ``slow thinking'' paradig…

DiversityLanguage ModelingLanguage ModellingMathematical Reasoning+1