paper-with-me

홈 › Papers

Automated curricula through setter-solver interactions

2019-09-27 · Sebastien Racaniere, Andrew K. Lampinen, Adam Santoro, David P. Reichert, Vlad Firoiu, Timothy P. Lillicrap

Reinforcement learning algorithms use correlations between policies and rewards to improve agent performance. But in dynamic or sparsely rewarding environments these correlations are often too small, or rewarding events are too infrequent to make learning feasible. Human education instead relies on curricula--the breakdown of tasks into simpler, static challenges with dense rewards--to build up to complex behaviors. While curricula are also useful for artificial agents, hand-crafting them is time consuming. This has lead researchers to explore automatic curriculum generation. Here we explore automatic curriculum generation in rich, dynamic environments. Using a setter-solver paradigm we show the importance of considering goal validity, goal feasibility, and goal coverage to construct useful curricula. We demonstrate the success of our approach in rich but sparsely rewarding 2D and 3D environments, where an agent is tasked to achieve a single goal selected from a set of possible goals that varies between episodes, and identify challenges for future work. Finally, we demonstrate the value of a novel technique that guides agents towards a desired goal distribution. Altogether, these results represent a substantial step towards applying automatic task curricula to learn complex, otherwise unlearnable goals, and to our knowledge are the first to demonstrate automated curriculum generation for goal-conditioned agents in environments where the possible goals vary between episodes.

📄 PDF Abstract BibTeX arXiv:1909.12892

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Automated curriculum generation through setter-solver interactions

2020-05-01 · ICLR 2020 1 · Sebastien Racaniere, Andrew Lampinen, Adam Santoro, David Reichert 외

Reinforcement learning algorithms use correlations between policies and rewards to improve agent performance. But in dynamic or sparsely rewarding environments these correlations are often too small, or rewarding even…

Verifier-Backed Hard Problem Generation for Mathematical Reasoning

2026-05-07 · Yuhang Lai, Jiazhan Feng, Yee Whye Teh, Ning Miao arxiv

Large Language Models (LLMs) demonstrate strong capabilities for solving scientific and mathematical problems, yet they struggle to produce valid, challenging, and novel problems - an essential component for advancing LL…

Mathematical Reasoning

Monopoly agenda control with privately informed voters

2024-02-09 · Kirill S. Evdokimov

An agenda-setter repeatedly proposes a spatial policy to voters until some proposal is accepted. Voters have distinct but correlated preferences and receive private signals about the common state. I investigate whether t…

Diversity

Who Controls the Agenda Controls the Polity

2022-12-02 · S. Nageeb Ali, B. Douglas Bernheim, Alexander W. Bloedel, Silvia Console Battilana

This paper models legislative decision-making with an agenda setter who can propose policies sequentially, tailoring each proposal to the status quo that prevails after prior votes. Voters are sophisticated and the agend…

Decision Making

Head-Internal Relatives in Japanese as Rich Context-Setters

2013-11-01 · PACLIC 2013 11 · Tohru Seraku