paper-with-me

홈 › Papers

Reward Gaming in Conditional Text Generation

2022-11-16 · Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, He He

To align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations. Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift. We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model. While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work.

📄 PDF Abstract BibTeX arXiv:2211.08714

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Text GenerationReinforcement Learning (RL)Text Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack

2024-10-09 · Leo McKee-Reid, Christoph Sträter, Maria Angelica Martinez, Joe Needham 외

Previous work has shown that training "helpful-only" LLMs with reinforcement learning on a curriculum of gameable environments can lead models to generalize to egregious specification gaming, such as editing their own re…

In-Context Reinforcement Learningreinforcement-learningReinforcement Learning

ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

2025-11-24 · Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie 외 arxiv

Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more d…

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

2024-06-14 · Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud 외

In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like syco…

Language ModellingLarge Language Model

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

2023-08-31 · ICCV 2023 1 · Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 외

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape…

3D Shape GenerationContrastive Learningcross-modal alignmentFace Generation+2

ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

2025-05-08 · Wanjiang Weng, Xiaofeng Tan, Hongsong Wang, Pan Zhou

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critic…

Motion Generation