paper-with-me

홈 › Papers

Self-supervised Sequential Information Bottleneck for Robust Exploration in Deep Reinforcement Learning

2022-09-12 · Bang You, Jingming Xie, Youping Chen, Jan Peters, Oleg Arenz

Effective exploration is critical for reinforcement learning agents in environments with sparse rewards or high-dimensional state-action spaces. Recent works based on state-visitation counts, curiosity and entropy-maximization generate intrinsic reward signals to motivate the agent to visit novel states for exploration. However, the agent can get distracted by perturbations to sensor inputs that contain novel but task-irrelevant information, e.g. due to sensor noise or changing background. In this work, we introduce the sequential information bottleneck objective for learning compressed and temporally coherent representations by modelling and compressing sequential predictive information in time-series observations. For efficient exploration in noisy environments, we further construct intrinsic rewards that capture task-relevant state novelty based on the learned representations. We derive a variational upper bound of our sequential information bottleneck objective for practical optimization and provide an information-theoretic interpretation of the derived upper bound. Our experiments on a set of challenging image-based simulated control tasks show that our method achieves better sample efficiency, and robustness to both white noise and natural video backgrounds compared to state-of-art methods based on curiosity, entropy maximization and information-gain.

📄 PDF Abstract BibTeX arXiv:2209.05333

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningEfficient Explorationreinforcement-learningReinforcement Learning (RL)Time SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

Dynamic Bottleneck for Robust Self-Supervised Exploration

2021-10-20 · NeurIPS 2021 12 · Chenjia Bai, Lingxiao Wang, Lei Han, Animesh Garg 외

Exploration methods based on pseudo-count of transitions or curiosity of dynamics have achieved promising results in solving reinforcement learning with sparse rewards. However, such methods are usually sensitive to envi…

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

2026-02-04 · Pingyue Zhang, Zihan Huang, Yue Wang, Jieyu Zhang 외 arxiv

Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploratio…

Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation

2022-03-06 · CVPR 2022 1 · Qi Chen, Lingxiao Yang, JianHuang Lai, Xiaohua Xie

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correla…

Semantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Self-Supervised Information Bottleneck for Deep Multi-View Subspace Clustering

2022-04-26 · Shiye Wang, Changsheng Li, Yanming Li, Ye Yuan 외

In this paper, we explore the problem of deep multi-view subspace clustering framework from an information-theoretic point of view. We extend the traditional information bottleneck principle to learn common information a…

ClusteringMulti-view Subspace Clustering

BottleSum: Unsupervised and Self-supervised Sentence Summarization using the Information Bottleneck Principle

2019-09-16 · IJCNLP 2019 11 · Peter West, Ari Holtzman, Jan Buys, Yejin Choi

The principle of the Information Bottleneck (Tishby et al. 1999) is to produce a summary of information X optimized to predict some other relevant information Y. In this paper, we propose a novel approach to unsupervised…

Abstractive Text SummarizationExtractive SummarizationLanguage ModelingLanguage Modelling+4