paper-with-me

홈 › Papers

Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

2026-07-14 · Amber Srivastava arxiv

Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs experienced by the agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem -- a formal expression for the gradient of the value function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{π,ξ}(s,a,ζ)$, which comprises two copies of the environment parameters: $ζ$ governs the cost and transition dynamics at the current state--action pair, while $ξ$ governs the future rollouts. This decoupling yields a tractable closed-form gradient expression and is essential to the theorem's derivation. Building on this result, we develop a model-free algorithm that simultaneously learns the optimal policy and the environment parameters. We demonstrate the efficacy of our framework on a UAV network design problem, where the optimal UAV placement (environment parameters) and communication routes (governed by the policy) are learned jointly to minimize the total communication cost in the network.

📄 PDF Abstract BibTeX arXiv:2607.12590

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Off-Policy Actor-Critic with Emphatic Weightings

2021-11-16 · Eric Graves, Ehsan Imani, Raksha Kumaraswamy, Martha White

A variety of theoretically-sound policy gradient algorithms exist for the on-policy setting due to the policy gradient theorem, which provides a simplified form for the gradient. The off-policy setting, however, has been…

The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations

2024-01-24 · Matthias Lehmann

In recent years, various powerful policy gradient algorithms have been proposed in deep reinforcement learning. While all these algorithms build on the Policy Gradient Theorem, the specific design choices differ signific…

continuous-controlContinuous ControlDeep Reinforcement LearningLearning Theory

Correcting discount-factor mismatch in on-policy policy gradient methods

2023-06-23 · Fengdi Che, Gautham Vasan, A. Rupam Mahmood

The policy gradient theorem gives a convenient form of the policy gradient in terms of three factors: an action value, a gradient of the action likelihood, and a state distribution involving discounting called the \emph{…

OpenAI GymPolicy Gradient Methods

Policy Gradient Methods in the Presence of Symmetries and State Abstractions

2023-05-09 · Prakash Panangaden, Sahand Rezaei-Shoshtari, Rosie Zhao, David Meger 외

Reinforcement learning (RL) on high-dimensional and complex problems relies on abstraction for improved efficiency and generalization. In this paper, we study abstraction in the continuous-control setting, and extend the…

continuous-controlContinuous ControlPolicy Gradient MethodsReinforcement Learning (RL)+1

Policy Gradient in Partially Observable Environments: Approximation and Convergence

2018-10-18 · Kamyar Azizzadenesheli, Yisong Yue, Animashree Anandkumar

Policy gradient is a generic and flexible reinforcement learning approach that generally enjoys simplicity in analysis, implementation, and deployment. In the last few decades, this approach has been extensively advanced…

Decision MakingPolicy Gradient MethodsReinforcement Learning