paper-with-me

홈 › Papers

Structure learning with Temporal Gaussian Mixture for model-based Reinforcement Learning

2024-11-18 · Théophile Champion, Marek Grześ, Howard Bowman

Model-based reinforcement learning refers to a set of approaches capable of sample-efficient decision making, which create an explicit model of the environment. This model can subsequently be used for learning optimal policies. In this paper, we propose a temporal Gaussian Mixture Model composed of a perception model and a transition model. The perception model extracts discrete (latent) states from continuous observations using a variational Gaussian mixture likelihood. Importantly, our model constantly monitors the collected data searching for new Gaussian components, i.e., the perception model performs a form of structure learning (Smith et al., 2020; Friston et al., 2018; Neacsu et al., 2022) as it learns the number of Gaussian components in the mixture. Additionally, the transition model learns the temporal transition between consecutive time steps by taking advantage of the Dirichlet-categorical conjugacy. Both the perception and transition models are able to forget part of the data points, while integrating the information they provide within the prior, which ensure fast variational inference. Finally, decision making is performed with a variant of Q-learning which is able to learn Q-values from beliefs over states. Empirically, we have demonstrated the model's ability to learn the structure of several mazes: the model discovered the number of states and the transition probabilities between these states. Moreover, using its learned Q-values, the agent was able to successfully navigate from the starting position to the maze's exit.

📄 PDF Abstract BibTeX arXiv:2411.11511

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingModel-based Reinforcement LearningNavigateQ-LearningVariational Inference

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding

2023-12-27 · Sunoh Kim, Jungchan Cho, Joonsang Yu, Youngjoon Yoo 외

In the weakly supervised temporal video grounding study, previous methods use predetermined single Gaussian proposals which lack the ability to express diverse events described by the sentence query. To enhance the expre…

SentenceTemporal Sentence GroundingvalidVideo Grounding

Gaussian Temporal Awareness Networks for Action Localization

2019-09-09 · CVPR 2019 6 · Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian 외

Temporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faste…

Action Localizationobject-detectionObject DetectionVideo Understanding

Online reinforcement learning via sparse Gaussian mixture model Q-functions

2025-09-18 · Minh Vu, Konstantinos Slavakis arxiv

This paper introduces a structured and interpretable online policy-iteration framework for reinforcement learning (RL), built around the novel class of sparse Gaussian mixture model Q-functions (S-GMM-QFs). Extending ear…

Reinforcement Learning

Temporal Gaussian Mixture Layer for Videos

2018-03-16 · ICLR 2019 5 · AJ Piergiovanni, Michael S. Ryoo

We introduce a new convolutional layer named the Temporal Gaussian Mixture (TGM) layer and present how it can be used to efficiently capture longer-term temporal information in continuous activity videos. The TGM layer i…

Action DetectionActivity Detection

Task-Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian Processes

2020-06-19 · NeurIPS 2020 12 · Mengdi Xu, Wenhao Ding, Jiacheng Zhu, Zuxin Liu 외

Continuously learning to solve unseen tasks with limited experience has been extensively pursued in meta-learning and continual learning, but with restricted assumptions such as accessible task distributions, independent…

Continual LearningDecision MakingGaussian ProcessesMeta-Learning+5