paper-with-me

홈 › Papers

When Waiting is not an Option : Learning Options with a Deliberation Cost

2017-09-14 · Jean Harb, Pierre-Luc Bacon, Martin Klissarov, Doina Precup

Recent work has shown that temporally extended actions (options) can be learned fully end-to-end as opposed to being specified in advance. While the problem of "how" to learn options is increasingly well understood, the question of "what" good options should be has remained elusive. We formulate our answer to what "good" options should be in the bounded rationality framework (Simon, 1957) through the notion of deliberation cost. We then derive practical gradient-based learning algorithms to implement this objective. Our results in the Arcade Learning Environment (ALE) show increased performance and interpretability.

📄 PDF Abstract BibTeX arXiv:1709.04571

Code (1)

jeanharb/a2oc_delib 공식 구현

Tasks

Atari Games

Similar Papers 제목 키워드 기반

Learnings Options End-to-End for Continuous Action Tasks

2017-11-30 · Martin Klissarov, Pierre-Luc Bacon, Jean Harb, Doina Precup

We present new results on learning temporally extended actions for continuoustasks, using the options framework (Suttonet al.[1999b], Precup [2000]). In orderto achieve this goal we work with the option-critic architectu…

MuJoCo

Temporally Extended Mixture-of-Experts Models

2026-04-22 · Zeyu Shen, Peter Henderson arxiv

Mixture-of-Experts models, now popular for scaling capacity at fixed inference speed, switch experts at nearly every token. Once a model outgrows available GPU memory, this churn can render optimizations like offloading …

Reinforcement LearningContinual Learning

Deliberation Networks and How to Train Them

2022-11-06 · Qingyun Dou, Mark Gales

Deliberation networks are a family of sequence-to-sequence models, which have achieved state-of-the-art performance in a wide range of tasks such as machine translation and speech synthesis. A deliberation network consis…

Machine TranslationSpeech Synthesis

Preserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation

2026-04-29 · Ariel Sela arxiv

Multi-agent deliberation systems using large language models (LLMs) are increasingly proposed for policy simulation, yet they suffer from artificial consensus: evaluator agents converge on the same option regardless of t…

Paternalism and Deliberation: An Experiment on Making Formal Rules

2025-01-01 · Max R. P. Grossmann

This paper studies the relationship between soft and hard paternalism by examining two kinds of restriction: a waiting period and a hard limit (cap) on risk-seeking behavior. Mandatory waiting periods have been institute…