paper-with-me

홈 › Papers

Shaped Rewards Bias Emergent Language

2021-09-29 · Brendon Boldt, Yonatan Bisk, David R Mortensen

One of the primary characteristics of emergent phenomena is that they are determined by the basic properties of the system whence they emerge as opposed to explicitly designed constraints. Reinforcement learning is often used to elicit such phenomena which specifically arise from the pressure to maximize reward. We distinguish two types of rewards. The first is the base reward which is motivated directly by the task being solved. The second is shaped rewards which are designed specifically to make the task easier to learn by introducing biases in the learning process. The inductive bias which reward shaping introduces is problematic for emergent language experimentation because it biases the object of study: the emergent language. The fact that shaped rewards are intentionally designed conflicts with the basic premise of emergent phenomena arising from basic principles. In this paper, we use a simple sender-receiver navigation game to demonstrate how reward shaping can 1) explicitly bias the semantics of the learned language, 2) significantly change the entropy of the learned communication, and 3) mask the potential effects of other environmental variables of interest.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive Bias

Similar Papers 제목 키워드 기반

Evolving AI Collectives to Enhance Human Diversity and Enable Self-Regulation

2024-02-19 · Shiyang Lai, Yujin Potter, Junsol Kim, Richard Zhuang 외

Large language model behavior is shaped by the language of those with whom they interact. This capacity and their increasing prevalence online portend that they will intentionally or unintentionally "program" one another…

DiversityLanguage ModelingLanguage ModellingLarge Language Model

Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL

2025-10-15 · Mahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths 외 arxiv

In this work, we take a first step toward elucidating the mechanisms behind emergent exploration in unsupervised reinforcement learning. We study Single-Goal Contrastive Reinforcement Learning (SGCRL), a self-supervised …

Reinforcement Learning

Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties

2025-10-22 · Philipp J. Schneider, Lin Tian, Marian-Andrei Rizoiu arxiv

Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily, reciprocity, and social validation -- and what memory and learning mechanisms en…

U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models

2024-10-02 · Tung-Yu Wu, Pei-Yu Lo

Large language models (LLMs) have been shown to exhibit emergent abilities in some downstream tasks, where performance seems to stagnate at first and then improve sharply and unpredictably with scale beyond a threshold. …

Searching for Structure: Investigating Emergent Communication with Large Language Models

2024-12-10 · Tom Kouwenhoven, Max Peeperkorn, Tessa Verhoef

Human languages have evolved to be structured through repeated language learning and use. These processes introduce biases that operate during language acquisition and shape linguistic systems toward communicative effici…

Language Acquisition