paper-with-me

Papers

Dual Approximation Policy Optimization

2024-10-02 · Zhihan Xiong, Maryam Fazel, Lin Xiao

We propose Dual Approximation Policy Optimization (DAPO), a framework that incorporates general function approximation into policy mirror descent methods. In contrast to the popular approach of using the $L_2$-norm to measure function approximation errors, DAPO uses the dual Bregman divergence induced by the mirror map for policy projection. This duality framework has both theoretical and practical implications: not only does it achieve fast linear convergence with general function approximation, but it also includes several well-known practical methods as special cases, immediately providing strong convergence guarantees.

📄 PDF Abstract BibTeX arXiv:2410.01249

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

DAPO Dialogue-Adaptive Pre-training Objective (DAPO) is a pre-training objective for dialogue adaptation, which is designed to measure qualities of dialogues from multiple…

Similar Papers 제목 키워드 기반

Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces

2026-03-10 · Ji Gao, Caleb Ju, Guanghui Lan, Zhaohui Tong arxiv

Policy Dual Averaging (PDA) offers a principled Policy Mirror Descent (PMD) framework that more naturally admits value function approximation than standard PMD, enabling the use of approximate advantage (or Q-) functions…

Reinforcement Learning

Variance Reduced Policy Evaluation with Smooth Function Approximation

2019-12-01 · NeurIPS 2019 12 · Hoi-To Wai, Mingyi Hong, Zhuoran Yang, Zhaoran Wang 외

Policy evaluation with smooth and nonlinear function approximation has shown great potential for reinforcement learning. Compared to linear function approxi- mation, it allows for using a richer class of approximation fu…

Reinforcement Learning

Residual Deep Reinforcement Learning for Inverter-based Volt-Var Control

2024-08-13 · Qiong Liu, Ye Guo, Lirong Deng, Haotian Liu 외

A residual deep reinforcement learning (RDRL) approach is proposed by integrating DRL with model-based optimization for inverter-based volt-var control in active distribution networks when the accurate power flow model i…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning

Proximal Point Imitation Learning

2022-09-22 · Luca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk 외

This work develops new algorithms with rigorous efficiency guarantees for infinite horizon imitation learning (IL) with linear function approximation without restrictive coherence assumptions. We begin with the minimax f…

Imitation Learning

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

2025-05-07 · Axel Friedrich Wolter, Tobias Sutter

We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation. Motivated by the challenge of designing algorithms that l…