paper-with-me

홈 › Papers

Learning Sampling Policy for Faster Derivative Free Optimization

2021-04-09 · Zhou Zhai, Bin Gu, Heng Huang

Zeroth-order (ZO, also known as derivative-free) methods, which estimate the gradient only by two function evaluations, have attracted much attention recently because of its broad applications in machine learning community. The two function evaluations are normally generated with random perturbations from standard Gaussian distribution. To speed up ZO methods, many methods, such as variance reduced stochastic ZO gradients and learning an adaptive Gaussian distribution, have recently been proposed to reduce the variances of ZO gradients. However, it is still an open problem whether there is a space to further improve the convergence of ZO methods. To explore this problem, in this paper, we propose a new reinforcement learning based ZO algorithm (ZO-RL) with learning the sampling policy for generating the perturbations in ZO optimization instead of using random sampling. To find the optimal policy, an actor-critic RL algorithm called deep deterministic policy gradient (DDPG) with two neural network function approximators is adopted. The learned sampling policy guides the perturbed points in the parameter space to estimate a more accurate ZO gradient. To the best of our knowledge, our ZO-RL is the first algorithm to learn the sampling policy using reinforcement learning for ZO optimization which is parallel to the existing methods. Especially, our ZO-RL can be combined with existing ZO algorithms that could further accelerate the algorithms. Experimental results for different ZO optimization problems show that our ZO-RL algorithm can effectively reduce the variances of ZO gradient by learning a sampling policy, and converge faster than existing ZO algorithms in different scenarios.

📄 PDF Abstract BibTeX arXiv:2104.04405

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Model-Free $μ$-Synthesis: A Nonsmooth Optimization Perspective

2024-02-18 · Darioush Keivan, Xingang Guo, Peter Seiler, Geir Dullerud 외

In this paper, we revisit model-free policy search on an important robust control benchmark, namely $\mu$-synthesis. In the general output-feedback setting, there do not exist convex formulations for this problem, and he…

model

Linear interpolation gives better gradients than Gaussian smoothing in derivative-free optimization

2019-05-29 · Albert S. Berahas, Liyuan Cao, Krzysztof Choromanski, Katya Scheinberg

In this paper, we consider derivative free optimization problems, where the objective function is smooth but is computed with some amount of noise, the function evaluations are expensive and no derivative information is …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems

2018-12-20 · Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru 외

We study derivative-free methods for policy optimization over the class of linear policies. We focus on characterizing the convergence rate of these methods when applied to linear-quadratic systems, and study various set…

Complexity of Derivative-Free Policy Optimization for Structured $\mathcal{H}_\infty$ Control

2023-09-21 · NeurIPS 2023 11

The applications of direct policy search in reinforcement learning and continuous control have received increasing attention. In this work, we present novel theoretical results on the complexity of derivative-free policy…

Dynamic Anisotropic Smoothing for Noisy Derivative-Free Optimization

2024-05-02 · Sam Reifenstein, Timothee Leleu, Yoshihisa Yamamoto

We propose a novel algorithm that extends the methods of ball smoothing and Gaussian smoothing for noisy derivative-free optimization by accounting for the heterogeneous curvature of the objective function. The algorithm…

Bayesian OptimizationCombinatorial Optimization