paper-with-me

홈 › Papers

Entropic Desired Dynamics for Intrinsic Control

2021-12-01 · NeurIPS 2021 12 · Steven Hansen, Guillaume Desjardins, Kate Baumli, David Warde-Farley, Nicolas Heess, Simon Osindero, Volodymyr Mnih

An agent might be said, informally, to have mastery of its environment when it has maximised the effective number of states it can reliably reach. In practice, this often means maximizing the number of latent codes that can be discriminated from future states under some short time horizon (e.g. \cite{eysenbach2018diversity}). By situating these latent codes in a globally consistent coordinate system, we show that agents can reliably reach more states in the long term while still optimizing a local objective. A simple instantiation of this idea, \textbf{E}ntropic \textbf{D}esired \textbf{D}ynamics for \textbf{I}ntrinsic \textbf{C}on\textbf{T}rol (EDDICT), assumes fixed additive latent dynamics, which results in tractable learning and an interpretable latent space. Compared to prior methods, EDDICT's globally consistent codes allow it to be far more exploratory, as demonstrated by improved state coverage and increased unsupervised performance on hard exploration games such as Montezuma's Revenge.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Montezuma's Revenge

Similar Papers 제목 키워드 기반

Entropic optimal transport beyond product reference couplings: the Gaussian case on Euclidean space

2025-07-02 · Paul Freulon, Nikitas Georgakis, Victor Panaretos arxiv

The Optimal Transport (OT) problem with squared Euclidean cost consists in finding a coupling between two input measures that maximizes correlation. Consequently, the optimal coupling is often singular with respect to th…

Minimum intrinsic dimension scaling for entropic optimal transport

2023-06-06 · Austin J. Stromme

Motivated by the manifold hypothesis, which states that data with a high extrinsic dimension may yet have a low intrinsic dimension, we develop refined statistical bounds for entropic optimal transport that are sensitive…

EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control

2025-11-19 · Kai Yang, Xin Xu, Yangkun Chen, Weijie Liu 외 arxiv

Long-term training of large language models (LLMs) requires maintaining stable exploration to prevent the model from collapsing into sub-optimal behaviors. Entropy is crucial in this context, as it controls exploration a…

Reinforcement Learning

Pattern formation using an intrinsic optimal control approach

2025-05-02 · TianHao Li, Yibei Li, Zhixin Liu, Xiaoming Hu

This paper investigates a pattern formation control problem for a multi-agent system modeled with given interaction topology, in which $m$ of the $n$ agents are chosen as leaders and consequently a control signal is adde…

Entropic Dynamics of Exchange Rates and Options

2019-08-18 · Mohammad Abedi, Daniel Bartolomeo

An Entropic Dynamics of exchange rates is laid down to model the dynamics of foreign exchange rates, FX, and European Options on FX. The main objective is to represent an alternative framework to model dynamics. Entropic…