Learning Open Domain Multi-hop Search Using Reinforcement Learning
We propose a method to teach an automated agent to learn how to search for multi-hop paths of relations between entities in an open domain. The method learns a policy for directing existing information retrieval and machine reading resources to focus on relevant regions of a corpus. The approach formulates the learning problem as a Markov decision process with a state representation that encodes the dynamics of the search process and a reward structure that minimizes the number of documents that must be processed while still finding multi-hop paths. We implement the method in an actor-critic reinforcement learning algorithm and evaluate it on a dataset of search problems derived from a subset of English Wikipedia. The algorithm finds a family of policies that succeeds in extracting the desired information while processing fewer documents compared to several baseline heuristic algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalReading Comprehensionreinforcement-learningReinforcement LearningReinforcement Learning (RL)RetrievalSimilar Papers 제목 키워드 기반
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via reinforcement learning with…
Reinforcement LearningFact CheckingSim-Env: Decoupling OpenAI Gym Environments from Simulation Models
Reinforcement learning (RL) is one of the most active fields of AI research. Despite the interest demonstrated by the research community in reinforcement learning, the development methodology still lags behind, with a se…
OpenAI Gymreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1Deep Reinforcement Learning Radio Control and Signal Detection with KeRLym, a Gym RL Agent
This paper presents research in progress investigating the viability and adaptation of reinforcement learning using deep neural network based function approximation for the task of radio control and signal detection in t…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)OR-Gym: A Reinforcement Learning Library for Operations Research Problems
Reinforcement learning (RL) has been widely applied to game-playing and surpassed the best human-level performance in many domains, yet there are few use-cases in industrial or commercial settings. We introduce OR-Gym, a…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)MMM: Multilingual Mutual Reinforcement Effect Mix Datasets & Test with Open-domain Information Extraction Large Language Models
The Mutual Reinforcement Effect (MRE) represents a promising avenue in information extraction and multitasking research. Nevertheless, its applicability has been constrained due to the exclusive availability of MRE mix d…
Language ModelingLanguage ModellingLarge Language Modelnamed-entity-recognition+5