RUST: Latent Neural Scene Representations from Unposed Imagery
Inferring the structure of 3D scenes from 2D observations is a fundamental challenge in computer vision. Recently popularized approaches based on neural scene representations have achieved tremendous impact and have been applied across a variety of applications. One of the major remaining challenges in this space is training a single model which can provide latent representations which effectively generalize beyond a single scene. Scene Representation Transformer (SRT) has shown promise in this direction, but scaling it to a larger set of diverse scenes is challenging and necessitates accurately posed ground truth data. To address this problem, we propose RUST (Really Unposed Scene representation Transformer), a pose-free approach to novel view synthesis trained on RGB images alone. Our main insight is that one can train a Pose Encoder that peeks at the target image and learns a latent pose embedding which is used by the decoder for view synthesis. We perform an empirical investigation into the learned latent pose structure and show that it allows meaningful test-time camera transformations and accurate explicit pose readouts. Perhaps surprisingly, RUST achieves similar quality as methods which have access to perfect camera pose, thereby unlocking the potential for large-scale training of amortized neural scene representations.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderNovel View SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ERUPT: Efficient Rendering with Unposed Patch Transformer
This work addresses the problem of novel view synthesis in diverse scenes from small collections of RGB images. We propose ERUPT (Efficient Rendering with Unposed Patch Transformer) a state-of-the-art scene reconstructio…
Image GenerationNeRFNovel View SynthesisScene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations
A classical problem in computer vision is to infer a 3D scene representation from few images that can be used to render novel views at interactive rates. Previous work focuses on reconstructing pre-defined 3D representat…
DecoderNovel View SynthesisSemantic SegmentationUnposed 3DGS Reconstruction with Probabilistic Procrustes Mapping
3D Gaussian Splatting (3DGS) has emerged as a core technique for 3D representation. Its effectiveness largely depends on precise camera poses and accurate point cloud initialization, which are often derived from pretrain…
Pose EstimationPoint CloudsLearning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed mu…
Representation LearningScene UnderstandingLearning Robust Multi-Scale Representation for Neural Radiance Fields from Unposed Images
We introduce an improved solution to the neural image-based rendering problem in computer vision. Given a set of images taken from a freely moving camera at train time, the proposed approach could synthesize a realistic …
Camera Pose EstimationDepth EstimationDepth PredictionGraph Neural Network+2