paper-with-me

홈 › Papers

DPO: A Differential and Pointwise Control Approach to Reinforcement Learning

2024-04-24 · Minh Nguyen, Chandrajit Bajaj

Reinforcement learning (RL) in continuous state-action spaces remains challenging in scientific computing due to poor sample efficiency and lack of pathwise physical consistency. We introduce Differential Reinforcement Learning (Differential RL), a novel framework that reformulates RL from a continuous-time control perspective via a differential dual formulation. This induces a Hamiltonian structure that embeds physics priors and ensures consistent trajectories without requiring explicit constraints. To implement Differential RL, we develop Differential Policy Optimization (DPO), a pointwise, stage-wise algorithm that refines local movement operators along the trajectory for improved sample efficiency and dynamic alignment. We establish pointwise convergence guarantees, a property not available in standard RL, and derive a competitive theoretical regret bound of $O(K^{5/6})$. Empirically, DPO outperforms standard RL baselines on representative scientific computing tasks, including surface modeling, grid control, and molecular dynamics, under low-data and physics-constrained conditions.

📄 PDF Abstract BibTeX arXiv:2404.15617

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarkingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

DPO 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Mollified Value Learning

2026-02-26 · Hrishikesh Viswanath, Juanwu Lu, S. Talha Bukhari, Mihir Chauhan 외 arxiv

Offline goal-conditioned reinforcement learning (GCRL) learns goal-reaching behaviors from static datasets, but accurate value estimation remains challenging under limited state-action coverage. Existing physics-informed…

Representation LearningReinforcement Learning

Physics-informed Gaussian Processes as Linear Model Predictive Controller

2024-12-02 · Jörn Tebbe, Andreas Besginow, Markus Lange-Hegermann

We introduce a novel algorithm for controlling linear time invariant systems in a tracking problem. The controller is based on a Gaussian Process (GP) whose realizations satisfy a system of linear ordinary differential e…

Bayesian InferenceGaussian ProcessesModel Predictive Control

Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration

2025-07-21 · Hee Jun Yang, Minjung Gim, Yeoneung Kim arxiv

We propose a mesh-free policy iteration framework that combines classical dynamic programming with physics-informed neural networks (PINNs) to solve high-dimensional, nonconvex Hamilton--Jacobi--Isaacs (HJI) equations ar…

Multi-agent Reinforcement Learning

Optimal Federated Learning for Nonparametric Regression with Heterogeneous Distributed Differential Privacy Constraints

2024-06-10 · T. Tony Cai, Abhinav Chakraborty, Lasse Vuursteen

This paper studies federated learning for nonparametric regression in the context of distributed samples across different servers, each adhering to distinct differential privacy constraints. The setting we consider is he…

Federated LearningPrivacy Preservingregression

MD-NOMAD: Mixture density nonlinear manifold decoder for emulating stochastic differential equations and uncertainty propagation

2024-04-24 · Akshay Thakur, Souvik Chakraborty

We propose a neural operator framework, termed mixture density nonlinear manifold decoder (MD-NOMAD), for stochastic simulators. Our approach leverages an amalgamation of the pointwise operator learning neural architectu…

DecoderOperator learning