paper-with-me

홈 › Papers

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control

2026-03-28 · Carles Domingo-Enrich, Jiequn Han arxiv

Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints. In this work, we revisit and generalize Adjoint Matching, a recently proposed SOC-based method for learning optimal controls, and place it on a rigorous footing by deriving it from the Stochastic Maximum Principle (SMP). We formulate a general Hamiltonian adjoint matching objective for SOC problems with control-dependent drift and diffusion and convex running costs, and show that its expected value has the same first variation as the original SOC objective. As a consequence, critical points satisfy the Hamilton--Jacobi--Bellman (HJB) stationarity conditions. In the important practical case of state- and control-independent diffusion, we recover the lean adjoint matching loss previously introduced, which avoids second-order terms and whose critical points coincide with the optimal control under mild uniqueness assumptions. Numerical experiments confirm that the extra terms it discards become necessary once the diffusion is state-dependent. Finally, we show that adjoint matching can be precisely interpreted as a continuous-time method of successive approximations induced by the SMP, yielding a practical and implementable alternative to classical SMP-based algorithms, which are obstructed by intractable martingale terms in the stochastic setting. These results are also of independent interest to the stochastic control community, providing new implementable objectives and a viable pathway for SMP-based iterations in stochastic problems.

📄 PDF Abstract BibTeX arXiv:2604.08580

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching

2026-06-21 · Serge Thilges, Onur Celik, Denis Blessing, Emiliyan Gospodinov 외 arxiv

Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online RL remains limited by the challenge of sc…

Reinforcement Learning

Functional Adjoint Sampler: Scalable Sampling on Infinite Dimensional Spaces

2025-11-09 · Byoungwoo Park, Juho Lee, Guan-Horng Liu arxiv

Learning-based methods for sampling from the Gibbs distribution in finite-dimensional spaces have progressed quickly, yet theory and algorithmic design for infinite-dimensional function spaces remain limited. This gap pe…

SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations

2025-02-04 · Grigory Bartosh, Dmitry Vetrov, Christian A. Naesseth

The Latent Stochastic Differential Equation (SDE) is a powerful tool for time series and sequence modeling. However, training Latent SDEs typically relies on adjoint sensitivity methods, which depend on simulation and ba…

SensitivityTime Series

Stochastic Operator Network: A Stochastic Maximum Principle Based Approach to Operator Learning

2025-07-10 · Ryan Bausback, Jingqiao Tang, Lu Lu, Feng Bao 외 arxiv

We develop a novel framework for uncertainty quantification in operator learning, the Stochastic Operator Network (SON). SON combines the stochastic optimal control concepts of the Stochastic Neural Network (SNN) with th…

Time-Reversal of Stochastic Maximum Principle

2024-03-04 · Amirhossein Taghvaei

Stochastic maximum principle (SMP) specifies a necessary condition for the solution of a stochastic optimal control problem. The condition involves a coupled system of forward and backward stochastic differential equatio…