Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition
We tackle the problem of object-centric learning on point clouds, which is crucial for high-level relational reasoning and scalable machine intelligence. In particular, we introduce a framework, SPAIR3D, to factorize a 3D point cloud into a spatial mixture model where each component corresponds to one object. To model the spatial mixture model on point clouds, we derive the Chamfer Mixture Loss, which fits naturally into our variational training pipeline. Moreover, we adopt an object-specification scheme that describes each object's location relative to its local voxel grid cell. Such a scheme allows SPAIR3D to model scenes with an arbitrary number of objects. We evaluate our method on the task of unsupervised scene decomposition. Experimental results demonstrate that SPAIR3D has strong scalability and is capable of detecting and segmenting an unknown number of objects from a point cloud in an unsupervised manner.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectRelational ReasoningSemantic SegmentationSimilar Papers 제목 키워드 기반
Variational Inference for Scalable 3D Object-centric Learning
We tackle the task of scalable unsupervised object-centric representation learning on 3D scenes. Existing approaches to object-centric representation learning show limitations in generalizing to larger scenes as their le…
NeRFObjectRepresentation LearningVariational InferenceSIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
To help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scen…
Instance SegmentationObjectSemantic SegmentationExploiting Spatial Invariance for Scalable Unsupervised Object Tracking
The ability to detect and track objects in the visual world is a crucial skill for any intelligent agent, as it is a necessary precursor to any object-level reasoning process. Moreover, it is important that agents learn …
ObjectObject TrackingSpatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better handled by spatially indicating what to move and where to place it. Ad…
Trajectory PredictionUnsupervised Discovery of Object-Centric Neural Fields
We study inferring 3D object-centric scene representations from a single image. While recent methods have shown potential in unsupervised 3D object discovery from simple synthetic images, they fail to generalize to real-…
ObjectObject DiscoverySemantic SegmentationSystematic Generalization+1