paper-with-me

Papers

PACER: A Fully Push-forward-based Distributional Reinforcement Learning Algorithm

2023-06-11 · Wensong Bai, Chao Zhang, Yichao Fu, Peilin Zhao, Hui Qian, Bin Dai

In this paper, we propose the first fully push-forward-based distributional reinforcement learning algorithm, named PACER, which consists of a distributional critic, a stochastic actor and a sample-based encourager. Specifically, the push-forward operator is leveraged in both the critic and actor to model the return distributions and stochastic policies respectively, enabling them with equal modeling capability and thus enhancing the synergetic performance. Since it is infeasible to obtain the density function of the push-forward policies, novel sample-based regularizers are integrated in the encourager to incentivize efficient exploration and alleviate the risk of trapping into local optima. Moreover, a sample-based stochastic utility value policy gradient is established for the push-forward policy update, which circumvents the explicit demand of the policy density function in existing REINFORCE-based stochastic policy gradient. As a result, PACER fully utilizes the modeling capability of the push-forward operator and is able to explore a broader class of the policy space, compared with limited policy classes used in existing distributional actor critic algorithms (i.e. Gaussians). We validate the critical role of each component in our algorithm with extensive empirical studies. Experimental results demonstrate the superiority of our algorithm over the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2306.06637

Code (0)

등록된 구현이 없습니다.

Tasks

Continuous ControlDistributional Reinforcement LearningEfficient Explorationreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

2025-04-02 · Kun Ouyang, Yuanxin Liu, HaoNing Wu, Yi Liu 외

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems pr…

MMESpatial ReasoningVideo MMEVideo Understanding

SpaceRipple: Lightweight Semantic Delivery for Mission-Oriented LEO Earth Observation Satellite Networks

2026-06-25 · Ziyi Yang, Hao Yuan, Yunxiang Yi, Wenbo Wang 외 arxiv

Earth observation satellite networks generate massive volumes of high-resolution imagery, whereas inter-satellite and downlink resources remain limited. In many time-sensitive missions, ground users require mission-relev…

Optimal number of spacers in CRISPR arrays

2017-05-30

We estimate the number of spacers in a CRISPR array of a bacterium which maximizes its protection against a viral attack. The optimality follows from a competition between two trends: too few distinct spacers make the ba…

Spacer: Towards Engineered Scientific Inspiration

2025-08-25 · Minhyeong Lee, Suyoung Hwang, Seunghyun Moon, Geonho Nah 외 arxiv

Recent advances in LLMs have made automated scientific research the next frontline in the path to artificial superintelligence. However, these systems are bound either to tasks of narrow scope or the limited creative cap…

Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face

2023-02-28 · Christopher Akiki, Odunayo Ogundepo, Aleksandra Piktus, Xinyu Zhang 외

We present Spacerini, a tool that integrates the Pyserini toolkit for reproducible information retrieval research with Hugging Face to enable the seamless construction and deployment of interactive search engines. Spacer…

Information RetrievalRetrieval