paper-with-me

Papers

APRIL: Active Preference-learning based Reinforcement Learning

2012-08-05 · Riad Akrour, Marc Schoenauer, Michèle Sebag

This paper focuses on reinforcement learning (RL) with limited prior knowledge. In the domain of swarm robotics for instance, the expert can hardly design a reward function or demonstrate the target behavior, forbidding the use of both standard RL and inverse reinforcement learning. Although with a limited expertise, the human expert is still often able to emit preferences and rank the agent demonstrations. Earlier work has presented an iterative preference-based RL framework: expert preferences are exploited to learn an approximate policy return, thus enabling the agent to achieve direct policy search. Iteratively, the agent selects a new candidate policy and demonstrates it; the expert ranks the new demonstration comparatively to the previous best one; the expert's ranking feedback enables the agent to refine the approximate policy return, and the process is iterated. In this paper, preference-based reinforcement learning is combined with active ranking in order to decrease the number of ranking queries to the expert needed to yield a satisfactory policy. Experiments on the mountain car and the cancer treatment testbeds witness that a couple of dozen rankings enable to learn a competent policy.

📄 PDF Abstract BibTeX arXiv:1208.0984

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Preference-based Interactive Multi-Document Summarisation

2019-06-07 · Yang Gao, Christian M. Meyer, Iryna Gurevych

Interactive NLP is a promising paradigm to close the gap between automatic NLP systems and the human upper bound. Preference-based interactive learning has been successfully applied, but the existing methods require seve…

Active Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

APRIL: Interactively Learning to Summarise by Combining Active Preference Learning and Reinforcement Learning

2018-08-29 · EMNLP 2018 10 · Yang Gao, Christian M. Meyer, Iryna Gurevych

We propose a method to perform automatic document summarisation without using reference summaries. Instead, our method interactively learns from users' preferences. The merit of preference-based interactive summarisation…

Active Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation

2025-09-23 · Yuzhen Zhou, Jiajun Li, Yusheng Su, Gowtham Ramesh 외 arxiv

Reinforcement learning (RL) has become a cornerstone in advancing large-scale pre-trained language models (LLMs). Successive generations, including GPT-o series, DeepSeek-R1, Kimi-K1.5, Grok 4, and GLM-4.5, have relied o…

Reinforcement Learning

Towards Abstractive Timeline Summarisation using Preference-based Reinforcement Learning

2022-11-14 · Yuxuan Ye, Edwin Simpson

This paper introduces a novel pipeline for summarising timelines of events reported by multiple news sources. Transformer-based models for abstractive summarisation generate coherent and concise summaries of long documen…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Boosting Feedback Efficiency of Interactive Reinforcement Learning by Adaptive Learning from Scores

2023-07-11 · Shukai Liu, Chenming Wu, Ying Li, Liangjun Zhang

Interactive reinforcement learning has shown promise in learning complex robotic tasks. However, the process can be human-intensive due to the requirement of a large amount of interactive feedback. This paper presents a …

reinforcement-learningReinforcement Learning