paper-with-me

홈 › Papers

Adaptive Dataset Sampling by Deep Policy Gradient

2021-01-01 · Jaerin Lee, Kyoung Mu Lee

Mini-batch SGD is a predominant optimization method in deep learning. Several works aim to improve naïve random dataset sampling, which appears typically in deep learning literature, with additional prior to allow faster and better performing optimization. This includes, but not limited to, importance sampling and curriculum learning. In this work, we propose an alternative way: we think of sampling as a trainable agent and let this external model learn to sample mini-batches of training set items based on the current status and recent history of the learned model. The resulting adaptive dataset sampler, named RLSampler, is a policy network implemented with simple recurrent neural networks trained by a policy gradient algorithm. We demonstrate RLSampler on image classification benchmarks with several different learner architectures and show consistent performance gain over the originally reported scores. Moreover, either a pre-sampled sequence of indices or a pre-trained RLSampler turns out to be more effective than naïve random sampling regardless of the network initialization and model architectures. Our analysis reveals the possible existence of a model-agnostic sample sequence that best represents the dataset under mini-batch SGD optimization framework.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classification

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies

2025-08-01 · Nicholas E. Corrado, Josiah P. Hanna arxiv

Independent on-policy policy gradient algorithms are widely used for multi-agent reinforcement learning (MARL) in cooperative and no-conflict games, but they are known to converge sub-optimally when each agent's individu…

Multi-agent Reinforcement Learning

On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling

2023-11-14 · Nicholas E. Corrado, Josiah P. Hanna

On-policy reinforcement learning (RL) algorithms perform policy updates using i.i.d. trajectories collected by the current policy. However, after observing only a finite number of trajectories, on-policy sampling may pro…

MuJoCoreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Experience Selection for Policy Gradient

2020-02-17 · Saad Mohamad, Giovanni Montana

Policy gradient reinforcement learning (RL) algorithms have achieved impressive performance in challenging learning tasks such as continuous control, but suffer from high sample complexity. Experience replay is a commonl…

continuous-controlContinuous ControlOpenAI GymReinforcement Learning+1

Momentum-Based Policy Gradient Methods

2020-07-13 · ICML 2020 1 · Feihu Huang, Shangqian Gao, Jian Pei, Heng Huang

In the paper, we propose a class of efficient momentum-based policy gradient methods for the model-free reinforcement learning, which use adaptive learning rates and do not require any large batches. Specifically, we pro…

Policy Gradient Methods

Learning Sampling Policy for Faster Derivative Free Optimization

2021-04-09 · Zhou Zhai, Bin Gu, Heng Huang

Zeroth-order (ZO, also known as derivative-free) methods, which estimate the gradient only by two function evaluations, have attracted much attention recently because of its broad applications in machine learning communi…

reinforcement-learningReinforcement Learning (RL)