paper-with-me

홈 › Papers

Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

2022-10-03 · Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, Yejin Choi

We tackle the problem of aligning pre-trained large language models (LMs) with human preferences. If we view text generation as a sequential decision-making problem, reinforcement learning (RL) appears to be a natural conceptual framework. However, using RL for LM-based generation faces empirical challenges, including training instability due to the combinatorial action space, as well as a lack of open-source libraries and benchmarks customized for LM alignment. Thus, a question rises in the research community: is RL a practical paradigm for NLP? To help answer this, we first introduce an open-source modular library, RL4LMs (Reinforcement Learning for Language Models), for optimizing language generators with RL. The library consists of on-policy RL algorithms that can be used to train any encoder or encoder-decoder LM in the HuggingFace library (Wolf et al. 2020) with an arbitrary reward function. Next, we present the GRUE (General Reinforced-language Understanding Evaluation) benchmark, a set of 6 language generation tasks which are supervised not by target strings, but by reward functions which capture automated measures of human preference.GRUE is the first leaderboard-style evaluation of RL algorithms for NLP tasks. Finally, we introduce an easy-to-use, performant RL algorithm, NLPO (Natural Language Policy Optimization)} that learns to effectively reduce the combinatorial action space in language generation. We show 1) that RL techniques are generally better than supervised methods at aligning LMs to human preferences; and 2) that NLPO exhibits greater stability and performance than previous policy gradient methods (e.g., PPO (Schulman et al. 2017)), based on both automatic and human evaluations.

📄 PDF Abstract BibTeX arXiv:2210.01241

Code (3)

allenai/rl4lms 공식 구현 pytorch
debjitpaul/refiner pytorch
tedmoskovitz/constrainedrl4lms pytorch

Tasks

Decision MakingPolicy Gradient Methodsreinforcement-learningReinforcement Learning (RL)Sequential Decision MakingText Generation

Methods 이 논문이 사용한 방법론

Library 설명 없음
Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…

Similar Papers 제목 키워드 기반

GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models

2025-09-15 · Min Zeng, Jingfei Sun, Xueyou Luo, Caiquan Liu 외 arxiv

In natural language processing tasks, pure reinforcement learning (RL) fine-tuning methods often suffer from inefficient exploration and slow convergence; while supervised fine-tuning (SFT) methods, although efficient in…

Reinforcement LearningText Classification

MLGym: A New Framework and Benchmark for Advancing AI Research Agents

2025-02-20 · Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov 외

We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research o…

Reinforcement Learning (RL)

GLGE: A New General Language Generation Evaluation Benchmark

2020-11-24 · Findings (ACL) 2021 8 · Dayiheng Liu, Yu Yan, Yeyun Gong, Weizhen Qi 외

Multi-task benchmarks such as GLUE and SuperGLUE have driven great progress of pretraining and transfer learning in Natural Language Processing (NLP). These benchmarks mostly focus on a range of Natural Language Understa…

Natural Language UnderstandingText GenerationTransfer Learning

Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks

2025-10-07 · Rushiv Arora arxiv

Multi-task reinforcement learning often relies on task metadata -- such as brief natural-language descriptions -- to guide behavior across diverse objectives. We present Lexical Policy Networks (LEXPOL), a language-condi…

Reinforcement Learning

Survey on reinforcement learning for language processing

2021-04-12 · Victor Uc-Cetina, Nicolas Navarro-Guerrero, Anabel Martin-Gonzalez, Cornelius Weber 외

In recent years some researchers have explored the use of reinforcement learning (RL) algorithms as key components in the solution of various natural language processing tasks. For instance, some of these algorithms leve…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Survey