paper-with-me

홈 › Papers

DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning

2025-02-28 · Pengcheng Jiang, Jiacheng Lin, Lang Cao, Runchu Tian, SeongKu Kang, Zifeng Wang, Jimeng Sun, Jiawei Han

Information retrieval systems are crucial for enabling effective access to large document collections. Recent approaches have leveraged Large Language Models (LLMs) to enhance retrieval performance through query augmentation, but often rely on expensive supervised learning or distillation techniques that require significant computational resources and hand-labeled data. We introduce DeepRetrieval, a reinforcement learning (RL) approach that trains LLMs for query generation through trial and error without supervised data (reference query). Using retrieval metrics as rewards, our system generates queries that maximize retrieval performance. DeepRetrieval outperforms leading methods on literature search with 65.07% (vs. previous SOTA 24.68%) recall for publication search and 63.18% (vs. previous SOTA 32.11%) recall for trial search using real-world search engines. DeepRetrieval also dominates in evidence-seeking retrieval, classic information retrieval and SQL database search. With only 3B parameters, it outperforms industry-leading models like GPT-4o and Claude-3.5-Sonnet on 11/13 datasets. These results demonstrate that our RL approach offers a more efficient and effective paradigm for information retrieval. Our data and code are available at: https://github.com/pat-jj/DeepRetrieval.

📄 PDF Abstract BibTeX arXiv:2503.00223

Code (1)

pat-jj/deepretrieval 공식 구현 pytorch

Tasks

Information Retrievalreinforcement-learningReinforcement LearningReinforcement Learning (RL)Retrieval

Similar Papers 제목 키워드 기반

Large Language Models are Built-in Autoregressive Search Engines

2023-05-16 · Noah Ziems, Wenhao Yu, Zhihan Zhang, Meng Jiang

Document retrieval is a key stage of standard Web search engines. Existing dual-encoder dense retrievers obtain representations for questions and documents independently, allowing for only shallow interactions between th…

Open-Domain Question AnsweringQuestion AnsweringRetrieval

Towards Competitive Search Relevance For Inference-Free Learned Sparse Retrievers

2024-11-07 · Zhichao Geng, Dongyu Ru, Yang Yang

Learned sparse retrieval, which can efficiently perform retrieval through mature inverted-index engines, has garnered growing attention in recent years. Particularly, the inference-free sparse retrievers are attractive a…

Knowledge DistillationRetrievalZero Shot on BEIR (Inference Free Model)

ReasonIR: Training Retrievers for Reasoning Tasks

2025-04-29 · Rulin Shao, Rui Qiao, Varsha Kishore, Niklas Muennighoff 외

We present ReasonIR-8B, the first retriever specifically trained for general reasoning tasks. Existing retrievers have shown limited gains on reasoning tasks, in part because existing training datasets focus on short fac…

Information RetrievalMMLURAGSynthetic Data Generation

LitSearch: A Retrieval Benchmark for Scientific Literature Search

2024-07-10 · Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal 외

Literature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?" pose significant challenges for modern search engines and retrieval systems. These questions often…

ArticlesRerankingRetrieval

CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning

2021-12-16 · Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter 외

Compared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dia…

Conversational Question AnsweringPassage RetrievalQuestion Answeringreinforcement-learning+3