paper-with-me

홈 › Papers

Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning

2026-02-26 · Chris Samarinas, Haw-Shiuan Chang, Hamed Zamani arxiv

Reinforcement learning has emerged as an effective paradigm for training large language models to interleave reasoning with search engine calls. However, existing approaches face a fundamental credit assignment problem: methods like Search-R1 assign a single outcome reward to the entire multi-step trajectory, providing no signal about which reasoning or retrieval decisions were responsible for success or failure. Process-reward methods such as StepSearch introduce step-level supervision but still sample complete trajectories independently, so advantage estimates at any given step are contaminated by the randomness of all other steps. We propose SLATE (Step-Level Advantage estimation for Truncated Exploration), which addresses both problems through two complementary ideas. First, truncated step-level sampling generates k continuations from a shared prefix, isolating all variation to a single decision point. We prove this reduces the variance of advantage estimates by up to a factor of T compared to full-trajectory sampling for T-step trajectories, the first formal variance guarantee for step-level RL in retrieval-augmented reasoning. Second, dense, decomposed process rewards separately evaluate reasoning quality, query quality, and answer correctness on a ternary scale via an LLM judge, providing richer supervision than binary outcome signals or heuristic step-level scores. Experiments on seven QA benchmarks show that SLATE consistently outperforms both sparse-reward and process-reward baselines, achieving a 7.0% relative improvement over Search-R1 on the 7B model and 30.7% on the 3B model. Gains are largest on challenging multi-hop tasks, and ablations confirm that truncated sampling and dense rewards provide complementary benefits.

📄 PDF Abstract BibTeX arXiv:2602.23440

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

2026-02-05 · Shuo Nie, Hexuan Deng, Chao Wang, Ruiyu Fang 외 arxiv

As large language models become smaller and more efficient, small reasoning models (SRMs) are crucial for enabling chain-of-thought (CoT) reasoning in resource-constrained settings. However, they are prone to faithfulnes…

Reinforcement Learning

TreeRPO: Tree Relative Policy Optimization

2025-06-05 · Zhicheng Yang, Zhijiang Guo, Yinya Huang, Xiaodan Liang 외

Large Language Models (LLMs) have shown remarkable reasoning capabilities through Reinforcement Learning with Verifiable Rewards (RLVR) methods. However, a key limitation of existing approaches is that rewards defined at…

Math

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

2026-05-28 · Yuchen Liu, Yingjie Feng, Lixiong Qin, Jiasi Chen 외 arxiv

In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward methods typically rely on costly tree sampling. We view world knowle…

Adaptive Multi-resolution Hash-Encoding Framework for INR-based Dental CBCT Reconstruction with Truncated FOV

2025-06-14 · Hyoung Suk Park, Kiwan Jeon

Implicit neural representation (INR), particularly in combination with hash encoding, has recently emerged as a promising approach for computed tomography (CT) image reconstruction. However, directly applying INR techniq…

Computational EfficiencyComputed Tomography (CT)Image Reconstruction

Nonparametric Bayesian Optimization for General Rewards

2026-02-07 · Zishi Zhang, Tao Ren, Yijie Peng arxiv

This work focuses on Bayesian optimization (BO) under reward model uncertainty. We propose the first BO algorithm that achieves no-regret guarantee in a general reward setting, requiring only Lipschitz continuity of the …