SPECTRA: Sparse Entity-centric Transitions
Learning an agent that interacts with objects is ubiquituous in many RL tasks. In most of them the agent's actions have sparse effects : only a small subset of objects in the visual scene will be affected by the action taken. We introduce SPECTRA, a model for learning slot-structured transitions from raw visual observations that embodies this sparsity assumption. Our model is composed of a perception module that decomposes the visual scene into a set of latent objects representations (i.e. slot-structured) and a transition module that predicts the next latent set slot-wise and in a sparse way. We show that learning a perception module jointly with a sparse slot-structured transition model not only biases the model towards more entity-centric perceptual groupings but also enables intrinsic exploration strategy that aims at maximizing the number of objects changed in the agent’s trajectory.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Entity Retrieval for Answering Entity-Centric Questions
The similarity between the question and indexed documents is a crucial factor in document retrieval for retrieval-augmented question answering. Although this is typically the only method for obtaining the relevant docume…
Entity RetrievalQuestion AnsweringRetrievalSparse Joint Transmission for Cell-Free Massive MIMO: A Sparse PCA Approach
Cell-free massive multiple-input multiple-output (MIMO) is a promising cellular network. In this network, a large number of distributed and multi-antenna access points (APs) jointly serve many single antenna users using …
CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation
Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together, these elements demand a level of fine-gr…
Video GenerationUnveiling the Significance of Toddler-Inspired Reward Transition in Goal-Oriented Reinforcement Learning
Toddlers evolve from free exploration with sparse feedback to exploiting prior experiences for goal-directed learning with denser rewards. Drawing inspiration from this Toddler-Inspired Reward Transition, we set out to e…
Reinforcement Learning (RL)NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative und…