Bootstrapping Top-down Information for Self-modulating Slot Attention
Object-centric learning (OCL) aims to learn representations of individual objects within visual scenes without manual supervision, facilitating efficient and effective visual reasoning. Traditional OCL methods primarily employ bottom-up approaches that aggregate homogeneous visual features to represent objects. However, in complex visual environments, these methods often fall short due to the heterogeneous nature of visual features within an object. To address this, we propose a novel OCL framework incorporating a top-down pathway. This pathway first bootstraps the semantics of individual objects and then modulates the model to prioritize features relevant to these semantics. By dynamically modulating the model based on its own output, our top-down pathway enhances the representational quality of objects. Our framework achieves state-of-the-art performance across multiple synthetic and real-world object-discovery benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectObject DiscoveryVisual ReasoningSimilar Papers 제목 키워드 기반
Bootstrapping NLU Models with Multi-task Learning
Bootstrapping natural language understanding (NLU) systems with minimal training data is a fundamental challenge of extending digital assistants like Alexa and Siri to a new language. A common approach that is adapted in…
General ClassificationMulti-Task LearningNatural Language UnderstandingWord EmbeddingsEffective Slot Filling via Weakly-Supervised Dual-Model Learning
Slot filling is a challenging task in Spoken Language Understanding (SLU). Supervised methods usually require large amounts of annotation to maintain desirable performance. A solution to relieve the heavy dependency on l…
slot-fillingSlot FillingSpoken Language UnderstandingSelf-Modulating Quantum Fast-Weight Programmers for Efficient Adaptive Sequential Learning
Recent advances in quantum machine learning have motivated efficient models for sequential data processing. In this paper, we propose Self-Modulating Quantum Fast Weight Programmers, or Self-Modulating QFWP, which extend…
Quantum Machine LearningSlotFM: A Motion Foundation Model with Slot Attention for Diverse Downstream Tasks
Wearable accelerometers are used for a wide range of applications, such as gesture recognition, gait analysis, and sports monitoring. Yet most existing foundation models focus primarily on classifying common daily activi…
Human Activity RecognitionGesture RecognitionTransfer-Free Data-Efficient Multilingual Slot Labeling
Slot labeling (SL) is a core component of task-oriented dialogue (ToD) systems, where slots and corresponding values are usually language-, task- and domain-specific. Therefore, extending the system to any new language-d…
Contrastive LearningCross-Lingual TransferSentencetoken-classification+1