paper-with-me

홈 › Papers

Real3D: Scaling Up Large Reconstruction Models with Real-World Images

2024-06-12 · Hanwen Jiang, QiXing Huang, Georgios Pavlakos

The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training procedure, they are hard to scale up beyond the existing datasets and they are not necessarily representative of the real distribution of object shapes. To address these limitations, in this paper, we introduce Real3D, the first LRM system that can be trained using single-view real-world images. Real3D introduces a novel self-training framework that can benefit from both the existing synthetic data and diverse single-view real images. We propose two unsupervised losses that allow us to supervise LRMs at the pixel- and semantic-level, even for training examples without ground-truth 3D or novel views. To further improve performance and scale up the image data, we develop an automatic data curation approach to collect high-quality examples from in-the-wild images. Our experiments show that Real3D consistently outperforms prior work in four diverse evaluation settings that include real and synthetic data, as well as both in-domain and out-of-domain shapes. Code and model can be found here: https://hwjiang1510.github.io/Real3D/

📄 PDF Abstract BibTeX arXiv:2406.08479

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

2024-12-18 · CVPR 2025 1 · Hanwen Jiang, Zexiang Xu, Desai Xie, Ziwen Chen 외

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real…

3D Reconstruction3D Scene Reconstruction

Robot Learning with Super-Linear Scaling

2024-12-02 · Marcel Torne, Arhan Jain, Jiayi Yuan, Vidaaranya Macha 외

Scaling robot learning requires data collection pipelines that scale favorably with human effort. In this work, we propose Crowdsourcing and Amortizing Human Effort for Real-to-Sim-to-Real(CASHER), a pipeline for scaling…

3D Reconstruction

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

2026-05-29 · Ulrich Prestel, Stefan Andreas Baumann, Nick Stracke, Björn Ommer arxiv

Self-supervised novel view synthesis (NVS) remains challenging to scale, despite the abundance of video data, largely due to the brittleness of training on realistic videos and the hard-to-predict scaling behavior of mul…

Novel View Synthesis

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

2026-04-12 · Chenhan Jiang, Yu Chen, Qingwen Zhang, Jifei Song 외 arxiv

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photo…

Novel View Synthesis

4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos

2025-06-09 · Zhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou 외

We propose 4DGT, a 4D Gaussian-based Transformer model for dynamic scene reconstruction, trained entirely on real-world monocular posed videos. Using 4D Gaussian as an inductive bias, 4DGT unifies static and dynamic comp…

Inductive Bias