paper-with-me

홈 › Papers

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

2026-05-27 · Yibo Zhao, Zichen Ding, Jiayi Wu, Zun Wang, Xiang Li arxiv

Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reasoning. However, the rapid growth of training methods has outpaced controlled comparison: existing works differ in retrieval corpora, reward designs, and training protocols, making it unclear what actually drives improvements. We present a controlled empirical study that isolates three under-explored dimensions of search agent training. First, we identify a critical data-coverage issue in the widely used Wikipedia 2018 corpus and show that correcting it alone yields larger gains than the differences between training algorithms. Second, we systematically compare outcome-based and process-based reward methods across three base models, finding that the simplest outcome-based approach achieves competitive or superior performance in most settings, and that process-level credit assignment can over-correct agent behavior. Third, we analyze training data diversity, off-policy data utilization, and search budget scaling, distilling practical guidelines for training effective search agents. Our code is available at https://github.com/YiboZhao624/SearchAgentReview.

📄 PDF Abstract BibTeX arXiv:2605.27881

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Matters to You? Towards Visual Representation Alignment for Robot Learning

2023-10-11 · Ran Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik 외

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs like RGB images, their rewards will inevitably use visual representa…

Zero-shot Generalization

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

2026-05-25 · Tianda Sun, Dimitar Kazakov arxiv

We test the standard RLVR tool-use recipe -- GRPO on Qwen2.5-7B-Instruct -- on a deliberately minimal knowledge-graph tool API: four Freebase navigation verbs over Complex WebQuestions. Under a self-verifiable retrieval …

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

2026-07-01 · Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini hf

In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matte…

Arithmetic Reasoning

ANN Search: Recall What Matters

2026-06-03 · Dimitris Dimitropoulos, Nikos Mamoulis arxiv

Approximate nearest neighbor (ANN) search has become a core primitive in information retrieval and modern machine learning tasks, from classification to retrieval-augmented generation. The community evaluates and tunes A…

Information RetrievalSemantic Similarity

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning

2026-04-25 · Xuanyue Zhong, Yuqiang Xie, Guanqun Bi, Jiangping Yang 외 arxiv

Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see \textit{what is happening} but fail to reason \textit{why it matters}. This semantic gap stems from the l…

Moment Retrieval