paper-with-me

홈 › Papers

REARANK: Reasoning Re-ranking Agent via Reinforcement Learning

2025-05-26 · Le Zhang, Bo wang, Xipeng Qiu, Siva Reddy, Aishwarya Agrawal

We present REARANK, a large language model (LLM)-based listwise reasoning reranking agent. REARANK explicitly reasons before reranking, significantly improving both performance and interpretability. Leveraging reinforcement learning and data augmentation, REARANK achieves substantial improvements over baseline models across popular information retrieval benchmarks, notably requiring only 179 annotated samples. Built on top of Qwen2.5-7B, our REARANK-7B demonstrates performance comparable to GPT-4 on both in-domain and out-of-domain benchmarks and even surpasses GPT-4 on reasoning-intensive BRIGHT benchmarks. These results underscore the effectiveness of our approach and highlight how reinforcement learning can enhance LLM reasoning capabilities in reranking.

📄 PDF Abstract BibTeX arXiv:2505.20046

Code (1)

lezhang7/rearank 공식 구현 pytorch

Tasks

Data AugmentationInformation RetrievalLanguage ModelingLanguage ModellingLarge Language Modelreinforcement-learningReinforcement LearningRerankingRe-RankingRetrieval

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search

2026-04-19 · Hansi Zeng, Liam Collins, Bhuvesh Kumar, Neil Shah 외 arxiv

Agentic search -- the task of training agents that iteratively reason, issue queries, and synthesize retrieved information to answer complex questions -- has achieved remarkable progress through reinforcement learning (R…

Reinforcement LearningDocument Ranking

PyVision-RL: Forging Open Agentic Vision Models via RL

2026-02-24 · Shitian Zhao, Shaoheng Lin, Ming Li, Haoquan Zhang 외 arxiv

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVi…

Reinforcement Learning

LegalMALR:Multi-Agent Query Understanding and LLM-Based Reranking for Chinese Statute Retrieval

2026-01-25 · Yunhan Li, Mingjie Xie, Gaoli Kang, Zihan Gong 외 arxiv

Statute retrieval is essential for legal assistance and judicial decision support, yet real-world legal queries are often implicit, multi-issue, and expressed in colloquial or underspecified forms. These characteristics …

Legal Reasoning

RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback

2025-10-05 · Tommy Mordo, Sagie Dekel, Omer Madmon, Moshe Tennenholtz 외 arxiv

Competitive search is a setting where document publishers modify them to improve their ranking in response to a query. Recently, publishers have increasingly leveraged LLMs to generate and modify competitive content. We …

Reinforcement Learning

GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation

2026-06-03 · Yuxiao Ye, Yiwen Zhang, Huiyuan Xie, Yuqin Huang 외 arxiv

LLM-based multi-agent systems are increasingly used for strategic decision-making tasks. In such settings, performance depends not only on individual model capabilities, but also on the policies by which agents interact …

Multi-agent Reinforcement Learning