paper-with-me

홈 › Papers

PAC-Bayesian Reinforcement Learning Trains Generalizable Policies

2025-10-12 · Abdelkrim Zitouni, Mehdi Hennequin, Juba Agoun, Ryan Horache, Nadia Kabachi, Omar Rivasplata arxiv

We derive a novel PAC-Bayesian generalization bound for reinforcement learning that explicitly accounts for Markov dependencies in the data, through the chain's mixing time. This contributes to overcoming challenges in obtaining generalization guarantees for reinforcement learning, where the sequential nature of data breaks the independence assumptions underlying classical bounds. The new bound provides non-vacuous certificates for modern off-policy algorithms such as Soft Actor-Critic. We demonstrate the practical utility of the bound through PB-SAC, a novel algorithm that optimizes the bound during training to guide exploration. Experiments across several continuous control tasks show that the proposed approach provides meaningful confidence certificates while maintaining competitive performance.

📄 PDF Abstract BibTeX arXiv:2510.10544

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContinuous Control

Similar Papers 제목 키워드 기반

Integrating Deep RL and Bayesian Inference for ObjectNav in Mobile Robotics

2026-03-26 · João Castelo-Branco, José Santos-Victor, Alexandre Bernardino arxiv

Autonomous object search is challenging for mobile robots operating in indoor environments due to partial observability, perceptual uncertainty, and the need to trade off exploration and navigation efficiency. Classical …

Reinforcement LearningBayesian Inference

Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization

2025-05-19 · Haochen Yuan, Minting Pan, Yunbo Wang, Siyu Gao 외

Reinforcement learning (RL) has shown significant promise for sequential portfolio optimization tasks, such as stock trading, where the objective is to maximize cumulative returns while minimizing risks using historical …

Offline RLPortfolio OptimizationReinforcement Learning (RL)Stock Prediction

Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics

2026-06-22 · Shujia Li, Jianshu Hu, Haiyu Zhang, Yunpeng Jiang 외 arxiv

Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current data-driven approaches primarily rely on motion capture datasets, whi…

Reinforcement Learning

L$^{2}$NAS: Learning to Optimize Neural Architectures via Continuous-Action Reinforcement Learning

2021-09-25 · Keith G. Mills, Fred X. Han, Mohammad Salameh, SEYED SAEED CHANGIZ REZAEI 외

Neural architecture search (NAS) has achieved remarkable results in deep neural network design. Differentiable architecture search converts the search over discrete architectures into a hyperparameter optimization proble…

Hyperparameter OptimizationNeural Architecture Searchreinforcement-learningReinforcement Learning (RL)

Bayesian Critique-Tune-Based Reinforcement Learning with Adaptive Pressure for Multi-Intersection Traffic Signal Control

2024-12-18 · Wenchang Duan, Zhenguo Gao, Jiwan He, Jinguo Xian

Adaptive Traffic Signal Control (ATSC) system is a critical component of intelligent transportation, with the capability to significantly alleviate urban traffic congestion. Although reinforcement learning (RL)-based met…

Bayesian Inferencereinforcement-learningReinforcement LearningReinforcement Learning (RL)+1