paper-with-me

홈 › Papers

StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

2025-05-21 · Ziliang Wang, Xuhui Zheng, Kang An, Cijun Ouyang, Jialu Cai, Yuhang Wang, Yichao Wu

Efficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively. Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document retrieval, achieving notable improvements in QA performance, but underperform on complex, multi-hop QA resulting from the sparse rewards from global signal only. To address this gap in existing research, we introduce StepSearch, a framework for search LLMs that trained with step-wise proximal policy optimization method. It consists of richer and more detailed intermediate search rewards and token-level process supervision based on information gain and redundancy penalties to better guide each search step. We constructed a fine-grained question-answering dataset containing sub-question-level search trajectories based on open source datasets through a set of data pipeline method. On standard multi-hop QA benchmarks, it significantly outperforms global-reward baselines, achieving 11.2% and 4.2% absolute improvements for 3B and 7B models over various search with RL baselines using only 19k training data, demonstrating the effectiveness of fine-grained, stepwise supervision in optimizing deep search LLMs. Our implementation is publicly available at https://github.com/zxh20001117/StepSearch.

📄 PDF Abstract BibTeX arXiv:2505.15107

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Igniting Language Intelligence: The Hitchhiker's Guide From Chain-of-Thought Reasoning to Language Agents

2023-11-20 · Zhuosheng Zhang, Yao Yao, Aston Zhang, Xiangru Tang 외

Large language models (LLMs) have dramatically enhanced the field of language intelligence, as demonstrably evidenced by their formidable empirical performance across a spectrum of complex reasoning tasks. Additionally, …

Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning

2026-02-26 · Chris Samarinas, Haw-Shiuan Chang, Hamed Zamani arxiv

Reinforcement learning has emerged as an effective paradigm for training large language models to interleave reasoning with search engine calls. However, existing approaches face a fundamental credit assignment problem: …

Reinforcement Learning

TikTok Search Recommendations: Governance and Research Challenges

2025-05-13 · Taylor Annabell, Robert Gorwa, Rebecca Scharlach, Jacob van de Kerkhof 외

Like other social media, TikTok is embracing its use as a search engine, developing search products to steer users to produce searchable content and engage in content discovery. Their recently developed product search re…

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards

2025-08-29 · Xiaolong Wei, Bo Lu, Xingyu Zhang, Zhejun Zhao 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable creative writing capabilities, yet their substantial computational demands hinder widespread use. Enhancing Small Language Models (SLMs) offers a promising altern…

Reinforcement Learning

DSBench: How Far Are Data Science Agents to Becoming Data Science Experts?

2024-09-12 · Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao 외

Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopp…