paper-with-me

홈 › Papers

RUST: Latent Neural Scene Representations from Unposed Imagery

2022-11-25 · CVPR 2023 1 · Mehdi S. M. Sajjadi, Aravindh Mahendran, Thomas Kipf, Etienne Pot, Daniel Duckworth, Mario Lucic, Klaus Greff

Inferring the structure of 3D scenes from 2D observations is a fundamental challenge in computer vision. Recently popularized approaches based on neural scene representations have achieved tremendous impact and have been applied across a variety of applications. One of the major remaining challenges in this space is training a single model which can provide latent representations which effectively generalize beyond a single scene. Scene Representation Transformer (SRT) has shown promise in this direction, but scaling it to a larger set of diverse scenes is challenging and necessitates accurately posed ground truth data. To address this problem, we propose RUST (Really Unposed Scene representation Transformer), a pose-free approach to novel view synthesis trained on RGB images alone. Our main insight is that one can train a Pose Encoder that peeks at the target image and learns a latent pose embedding which is used by the decoder for view synthesis. We perform an empirical investigation into the learned latent pose structure and show that it allows meaningful test-time camera transformations and accurate explicit pose readouts. Perhaps surprisingly, RUST achieves similar quality as methods which have access to perfect camera pose, thereby unlocking the potential for large-scale training of amortized neural scene representations.

📄 PDF Abstract BibTeX arXiv:2211.14306

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderNovel View Synthesis

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Adam 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

ERUPT: Efficient Rendering with Unposed Patch Transformer

2025-03-31 · CVPR 2025 1 · Maxim V. Shugaev, Vincent Chen, Maxim Karrenbach, Kyle Ashley 외

This work addresses the problem of novel view synthesis in diverse scenes from small collections of RGB images. We propose ERUPT (Efficient Rendering with Unposed Patch Transformer) a state-of-the-art scene reconstructio…

Image GenerationNeRFNovel View Synthesis

Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations

2021-11-25 · CVPR 2022 1 · Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann 외

A classical problem in computer vision is to infer a 3D scene representation from few images that can be used to render novel views at interactive rates. Previous work focuses on reconstructing pre-defined 3D representat…

DecoderNovel View SynthesisSemantic Segmentation

Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping

2025-07-24 · Chong Cheng, Zijian Wang, Sicheng Yu, Yu Hu 외 arxiv

3D Gaussian Splatting (3DGS) has emerged as a core technique for 3D representation. Its effectiveness largely depends on precise camera poses and accurate point cloud initialization, which are often derived from pretrain…

Pose EstimationPoint Clouds

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

2026-04-12 · Bo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu 외 arxiv

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed mu…

Representation LearningScene Understanding

Learning Robust Multi-Scale Representation for Neural Radiance Fields from Unposed Images

2023-11-08 · Nishant Jain, Suryansh Kumar, Luc van Gool

We introduce an improved solution to the neural image-based rendering problem in computer vision. Given a set of images taken from a freely moving camera at train time, the proposed approach could synthesize a realistic …

Camera Pose EstimationDepth EstimationDepth PredictionGraph Neural Network+2