paper-with-me

홈 › Papers

Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control

2025-10-20 · Chengxiu Hua, Jiawen Gu, Yushun Tang arxiv

Reinforcement learning (RL) has achieved significant success across a wide range of domains, however, most existing methods are formulated in discrete time. In this work, we introduce a novel RL method for continuous-time control, where stochastic differential equations govern state-action dynamics. Departing from traditional value function-based approaches, our key contribution is the characterization of continuous-time Q-functions via a martingale condition and the linking of diffusion policy scores to the action gradient of a learned continuous Q-function by the dynamic programming principle. This insight motivates Continuous Q-Score Matching (CQSM), a score-based policy improvement algorithm. Notably, our method addresses a long-standing challenge in continuous-time RL: preserving the action-evaluation capability of Q-functions without relying on time discretization. We further provide theoretical closed-form solutions for linear-quadratic (LQ) control problems within our framework. Numerical results in simulated environments demonstrate the effectiveness of our proposed method and compare it to popular baselines.

📄 PDF Abstract BibTeX arXiv:2510.17122

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

2023-12-18 · Michael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi Ma

Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning. This is due to their natural ability to optimize an expressive class of distributions …

Denoisingreinforcement-learningReinforcement Learning

Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement

2026-05-14 · Injin Kong, Hyoungjoon Lee, Yohan Jo arxiv

Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language denoising and token recovery. We propose DiHAL, a geometry-guided diffu…

Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning

2024-09-12 · Hanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao 외

Reinforcement Learning from human feedback (RLHF) has been shown a promising direction for aligning generative models with human intent and has also been explored in recent works for alignment of diffusion generative mod…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

2025-02-03 · Hanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao 외

Reinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models. Most works in this area use a discrete-time formul…

Energy-Weighted Flow Matching for Offline Reinforcement Learning

2025-03-06 · Shiyuan Zhang, Weitong Zhang, Quanquan Gu

This paper investigates energy guidance in generative modeling, where the target distribution is defined as $q(\mathbf x) \propto p(\mathbf x)\exp(-\beta \mathcal E(\mathbf x))$, with $p(\mathbf x)$ being the data distri…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)