paper-with-me

홈 › Papers

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

2024-12-18 · CVPR 2025 1 · Hanwen Jiang, Zexiang Xu, Desai Xie, Ziwen Chen, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang, Sai Bi, Xin Sun, Jiuxiang Gu, QiXing Huang, Georgios Pavlakos, Hao Tan

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth's properties for enhancing model capability, training stability, and generalization.

📄 PDF Abstract BibTeX arXiv:2412.14166

Code (0)

등록된 구현이 없습니다.

Tasks

3D Reconstruction3D Scene Reconstruction

Similar Papers 제목 키워드 기반

ISNN: Impact Sound Neural Network for Audio-Visual Object Classification

2018-09-01 · ECCV 2018 9 · Auston Sterling, Justin Wilson, Sam Lowe, Ming C. Lin

3D object geometry reconstruction remains a challenge when working with transparent, occluded, or highly reflective surfaces. While recent methods classify shape features using raw audio, we present a multimodal neural n…

ClassificationGeneral ClassificationObject

DeRainGS: Gaussian Splatting for Enhanced Scene Reconstruction in Rainy Environments

2024-08-21 · Shuhong Liu, Xiang Chen, Hongming Chen, Quanfeng Xu 외

Reconstruction under adverse rainy conditions poses significant challenges due to reduced visibility and the distortion of visual perception. These conditions can severely impair the quality of geometric maps, which is e…

3DGS3D Reconstruction

Multi-view reconstruction of bullet time effect based on improved NSFF model

2023-04-01 · Linquan Yu, Yan Gao, Yangtian Yan, Wentao Zeng

Bullet time is a type of visual effect commonly used in film, television and games that makes time seem to slow down or stop while still preserving dynamic details in the scene. It usually requires multiple sets of camer…

Neural RenderingOptical Flow Estimation

GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

2026-05-22 · Katharina Schmid, Nicolas von Lützow, Jozef Hladký, Angela Dai 외 arxiv

We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D genera…

3D Generation

City-Mesh3R: Simulation-Ready City-Scale 3D Mesh Reconstruction from Multi-View Images

2026-05-28 · Sayan Paul, Sourav Ghosh, Siddharth Katageri, Soumyadip Maity 외 arxiv

City-scale 3D surface reconstruction from multiview images for downstream 3D simulation, poses highly challenging problems due to the scale and complexity of urban scenes. Existing city-scale 3D reconstruction methods ba…

3D ReconstructionImage Clustering