paper-with-me

홈 › Papers

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

2023-12-18 · Michael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi Ma

Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning. This is due to their natural ability to optimize an expressive class of distributions over a continuous space. However, previous works fail to exploit the score-based structure of diffusion models, and instead utilize a simple behavior cloning term to train the actor, limiting their ability in the actor-critic setting. In this paper, we present a theoretical framework linking the structure of diffusion model policies to a learned Q-function, by linking the structure between the score of the policy to the action gradient of the Q-function. We focus on off-policy reinforcement learning and propose a new policy update method from this theory, which we denote Q-score matching. Notably, this algorithm only needs to differentiate through the denoising model rather than the entire diffusion model evaluation, and converged policies through Q-score matching are implicitly multi-modal and explorative in continuous domains. We conduct experiments in simulated environments to demonstrate the viability of our proposed method and compare to popular baselines. Source code is available from the project website: https://michaelpsenka.io/qsm.

📄 PDF Abstract BibTeX arXiv:2312.11752

Code (1)

Alescontrela/score_matching_rl 공식 구현 jax

Tasks

Denoisingreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

2025-09-29 · Shuchen Xue, Chongjian Ge, Shilong Zhang, Yichen Li 외 arxiv

Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where pre-training and RL post-training share the same log-likelihood formulation. In contrast, recent RL approach…

Reinforcement Learning

X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching

2026-07-30 · Tianyu Yang, Yiming Zeng, Wenzhe Cai, Yuqiang Yang 외 arxiv

Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's genera…

Reinforcement LearningVisual Navigation

Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation

2024-12-12 · Bofang Jia, Pengxiang Ding, Can Cui, Mingyang Sun 외

Visual-motor policy learning has advanced with architectures like diffusion-based policies, known for modeling complex robotic trajectories. However, their prolonged inference times hinder high-frequency control tasks re…

Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods

2025-02-03 · Oussama Zekri, Nicolas Boullé

Discrete diffusion models have recently gained significant attention due to their ability to process complex discrete structures for language modeling. However, fine-tuning these models with policy gradient methods, as i…

Language ModelingLanguage ModellingPolicy Gradient Methods

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

2026-01-02 · Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh, Anas El Houssaini 외 arxiv

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stochastic differential equation (SDE). Howe…

Continuous ControlImage Generation