Learning Global Spatial Information for Multi-View Object-Centric Models
Recently, several studies have been working on multi-view object-centric models, which predict unobserved views of a scene and infer object-centric representations from several observation views. In general, multi-object scenes can be uniquely determined if both the properties of individual objects and the spatial arrangement of objects are specified; however, existing multi-view object-centric models only infer object-level representations and lack spatial information. This insufficient modeling can degrade novel-view synthesis quality and make it difficult to generate novel scenes. We can model both spatial information and object representations by introducing hierarchical probabilistic model, which contains a global latent variable on top of object-level latent variables. However, how to execute inference and training with that hierarchical multi-view object-centric model is unclear. Therefore, we introduce several crucial components which help inference and training with the proposed model. We show that the proposed method achieves good inference quality and can also generate novel scenes.
Code (0)
등록된 구현이 없습니다.
Tasks
Novel View SynthesisObjectSimilar Papers 제목 키워드 기반
View N-gram Network for 3D Object Retrieval
How to aggregate multi-view representations of a 3D object into an informative and discriminative one remains a key challenge for multi-view 3D object retrieval. Existing methods either use view-wise pooling strategies w…
3D Object Retrieval3D Shape Classification3D Shape RetrievalObject+13M3D: Multi-view, Multi-path, Multi-representation for 3D Object Detection
3D visual perception tasks based on multi-camera images are essential for autonomous driving systems. Latest work in this field performs 3D object detection by leveraging multi-view images as an input and iteratively enh…
3D Object DetectionAutonomous DrivingObjectobject-detection+1TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive…
Spatial ReasoningMulti-view analysis of unregistered medical images using cross-view transformers
Multi-view medical image analysis often depends on the combination of information from multiple views. However, differences in perspective or other forms of misalignment can make it difficult to combine views effectively…
Medical Image AnalysisVISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial Attention
Detecting objects from LiDAR point clouds is of tremendous significance in autonomous driving. In spite of good progress, accurate and reliable 3D detection is yet to be achieved due to the sparsity and irregularity of L…
3D Object DetectionAutonomous Drivingobject-detectionObject Detection