paper-with-me

홈 › Papers

Bisimulation Metrics are Optimal Transport Distances, and Can be Computed Efficiently

2024-06-06 · Sergio Calo, Anders Jonsson, Gergely Neu, Ludovic Schwartz, Javier Segovia

We propose a new framework for formulating optimal transport distances between Markov chains. Previously known formulations studied couplings between the entire joint distribution induced by the chains, and derived solutions via a reduction to dynamic programming (DP) in an appropriately defined Markov decision process. This formulation has, however, not led to particularly efficient algorithms so far, since computing the associated DP operators requires fully solving a static optimal transport problem, and these operators need to be applied numerous times during the overall optimization process. In this work, we develop an alternative perspective by considering couplings between a flattened version of the joint distributions that we call discounted occupancy couplings, and show that calculating optimal transport distances in the full space of joint distributions can be equivalently formulated as solving a linear program (LP) in this reduced space. This LP formulation allows us to port several algorithmic ideas from other areas of optimal transport theory. In particular, our formulation makes it possible to introduce an appropriate notion of entropy regularization into the optimization problem, which in turn enables us to directly calculate optimal transport distances via a Sinkhorn-like method we call Sinkhorn Value Iteration (SVI). We show both theoretically and empirically that this method converges quickly to an optimal coupling, essentially at the same computational cost of running vanilla Sinkhorn in each pair of states. Along the way, we point out that our optimal transport distance exactly matches the common notion of bisimulation metrics between Markov chains, and thus our results also apply to computing such metrics, and in fact our algorithm turns out to be significantly more efficient than the best known methods developed so far for this purpose.

📄 PDF Abstract BibTeX arXiv:2406.04056

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음

Similar Papers 제목 키워드 기반

Approximate Policy Iteration with Bisimulation Metrics

2022-02-06 · Mete Kemertas, Allan Jepson

Bisimulation metrics define a distance measure between states of a Markov decision process (MDP) based on a comparison of reward sequences. Due to this property they provide theoretical guarantees in value function appro…

Continuous ControlRepresentation Learning

Sinkhorn Distances: Lightspeed Computation of Optimal Transportation Distances

2013-06-04 · NeurIPS 2013 · Marco Cuturi

Optimal transportation distances are a fundamental family of parameterized distances for histograms. Despite their appealing theoretical properties, excellent performance in retrieval tasks and intuitive formulation, the…

Retrieval

Sinkhorn Distances: Lightspeed Computation of Optimal Transport

2013-12-01 · NeurIPS 2013 12 · Marco Cuturi

Optimal transportation distances are a fundamental family of parameterized distances for histograms in the probability simplex. Despite their appealing theoretical properties, excellent performance and intuitive formulat…

Invariant Representations for Reinforcement Learning without Reconstruction

2021-01-01 · ICLR 2021 1 · Amy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal 외

We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations …

Causal InferenceMuJoCoreinforcement-learningReinforcement Learning+2

Learning Invariant Representations for Reinforcement Learning without Reconstruction

2020-06-18 · Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal 외

We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations …

Causal InferenceMuJoCoreinforcement-learningReinforcement Learning+2