paper-with-me

홈 › Papers

LVT: Large-Scale Scene Reconstruction via Local View Transformers

2025-09-29 · Tooba Imtiaz, Lucy Chai, Kathryn Heal, Xuan Luo, Jungyeon Park, Jennifer Dy, John Flynn arxiv

Large transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer's well-known quadratic complexity makes it difficult to scale these methods to large scenes. To address this challenge, we propose the Local View Transformer (LVT), a large-scale scene reconstruction and novel view synthesis architecture that circumvents the need for the quadratic attention operation. Motivated by the insight that spatially nearby views provide more useful signal about the local scene composition than distant views, our model processes all information in a local neighborhood around each view. To attend to tokens in nearby views, we leverage a novel positional encoding that conditions on the relative geometric transformation between the query and nearby views. We decode the output of our model into a 3D Gaussian Splat scene representation that includes both color and opacity view-dependence. Taken together, the Local View Transformer enables reconstruction of arbitrarily large, high-resolution scenes in a single forward pass. See our project page for results and interactive demos https://toobaimt.github.io/lvt/.

📄 PDF Abstract BibTeX arXiv:2509.25001

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View Synthesis

Similar Papers 제목 키워드 기반

Holistic Large-Scale Scene Reconstruction via Mixed Gaussian Splatting

2025-05-29 · Chuandong Liu, Huijiao Wang, Lei Yu, Gui-Song Xia

Recent advances in 3D Gaussian Splatting have shown remarkable potential for novel view synthesis. However, most existing large-scale scene reconstruction methods rely on the divide-and-conquer paradigm, which often lead…

3D Scene ReconstructionGPUNovel View Synthesis

Empowering Feed-Forward Reconstruction Models with Metric Scale via Satellite Images

2026-06-06 · Xianghui Ze, Yongjian Luo, Mengjun Chao, Zhenbo Song 외 arxiv

Feed-forward 3D reconstruction models have recently shown strong generalization across diverse scenes, yet most of them recover geometry only up to an unknown global scale. This scale ambiguity limits their use in applic…

Camera Localization3D ReconstructionDepth Estimation

GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

2026-05-22 · Katharina Schmid, Nicolas von Lützow, Jozef Hladký, Angela Dai 외 arxiv

We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D genera…

3D Generation

LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling

2025-07-03 · Jiahao Wu, Rui Peng, Jianbo Jiao, Jiayu Yang 외 arxiv

Due to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian sp…

Incremental Joint Learning of Depth, Pose and Implicit Scene Representation on Monocular Camera in Large-scale Scenes

2024-04-09 · Tianchen Deng, Nailin Wang, Chongdi Wang, Shenghai Yuan 외

Dense scene reconstruction for photo-realistic view synthesis has various applications, such as VR/AR, autonomous vehicles. However, most existing methods have difficulties in large-scale scenes due to three core challen…

Autonomous VehiclesDepth EstimationPose Estimation