When Waiting is not an Option : Learning Options with a Deliberation Cost
Recent work has shown that temporally extended actions (options) can be learned fully end-to-end as opposed to being specified in advance. While the problem of "how" to learn options is increasingly well understood, the question of "what" good options should be has remained elusive. We formulate our answer to what "good" options should be in the bounded rationality framework (Simon, 1957) through the notion of deliberation cost. We then derive practical gradient-based learning algorithms to implement this objective. Our results in the Arcade Learning Environment (ALE) show increased performance and interpretability.
Code (1)
Tasks
Atari GamesSimilar Papers 제목 키워드 기반
Learnings Options End-to-End for Continuous Action Tasks
We present new results on learning temporally extended actions for continuoustasks, using the options framework (Suttonet al.[1999b], Precup [2000]). In orderto achieve this goal we work with the option-critic architectu…
MuJoCoTemporally Extended Mixture-of-Experts Models
Mixture-of-Experts models, now popular for scaling capacity at fixed inference speed, switch experts at nearly every token. Once a model outgrows available GPU memory, this churn can render optimizations like offloading …
Reinforcement LearningContinual LearningDeliberation Networks and How to Train Them
Deliberation networks are a family of sequence-to-sequence models, which have achieved state-of-the-art performance in a wide range of tasks such as machine translation and speech synthesis. A deliberation network consis…
Machine TranslationSpeech SynthesisPreserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation
Multi-agent deliberation systems using large language models (LLMs) are increasingly proposed for policy simulation, yet they suffer from artificial consensus: evaluator agents converge on the same option regardless of t…
Paternalism and Deliberation: An Experiment on Making Formal Rules
This paper studies the relationship between soft and hard paternalism by examining two kinds of restriction: a waiting period and a hard limit (cap) on risk-seeking behavior. Mandatory waiting periods have been institute…