paper-with-me

Papers

Iterative Amortized Policy Optimization

2020-10-20 · NeurIPS 2021 12 · Joseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong Yue

Policy networks are a central feature of deep reinforcement learning (RL) algorithms for continuous control, enabling the estimation and sampling of high-value actions. From the variational inference perspective on RL, policy networks, when employed with entropy or KL regularization, are a form of amortized optimization, optimizing network parameters rather than the policy distributions directly. However, this direct amortized mapping can empirically yield suboptimal policy estimates. Given this perspective, we consider the more flexible class of iterative amortized optimizers. We demonstrate that the resulting technique, iterative amortized policy optimization, yields performance improvements over conventional direct amortization methods on benchmark continuous control tasks.

📄 PDF Abstract BibTeX arXiv:2010.10670

Code (1)

joelouismarino/variational_rl 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlDeep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Variational Inference

Similar Papers 제목 키워드 기반

Amortized Neural Optimization for Pre-Layout Signal Integrity Design Space Exploration using Differentiable Surrogates

2026-06-05 · Julian Withöft, Werner John, Emre Ecik, Ralf Brüning 외 arxiv

Pre-layout design space exploration (DSE) for high-speed signal integrity (SI) analysis is often limited by the computational cost of simulations and iterative optimization algorithms within modern electronic design auto…

Bayesian Program Learning by Decompiling Amortized Knowledge

2023-06-13 · Alessandro B. Palmarini, Christopher G. Lucas, N. Siddharth

DreamCoder is an inductive program synthesis system that, whilst solving problems, learns to simplify search in an iterative wake-sleep procedure. The cost of search is amortized by training a neural search policy, reduc…

Program inductionProgram Synthesis

Amortized Molecular Optimization via Group Relative Policy Optimization

2026-02-12 · Muhammad bin Javaid, Hasham Hussain, Ashima Khanna, Berke Kisin 외 arxiv

In structurally constrained molecular optimization, state-of-the-art methods restart an expensive oracle-driven search from scratch for every new input structure, scaling poorly to settings with many starting structures …

Reinforcement Learning

Amortized Projection Optimization for Sliced Wasserstein Generative Models

2022-03-25 · Khai Nguyen, Nhat Ho

Seeking informative projecting directions has been an important task in utilizing sliced Wasserstein distance in applications. However, finding these directions usually requires an iterative optimization procedure over t…

Cascaded Transformer for Robust and Scalable SLA Decomposition via Amortized Optimization

2026-01-17 · Cyril Shih-Huan Hsu arxiv

The evolution toward 6G networks increasingly relies on network slicing to provide tailored, End-to-End (E2E) logical networks over shared physical infrastructures. A critical challenge is effectively decomposing E2E Ser…