paper-with-me

홈 › Papers

VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling

2025-03-20 · Hyojun Go, Byeongjun Park, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, Changick Kim

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of real-world scenes, while ensuring generalization to arbitrary text prompts, previous methods fine-tune 2D generative models to jointly model camera poses and multi-view images. However, these methods suffer from instability when extending 2D generative models to joint modeling due to the modality gap, which necessitates additional models to stabilize training and inference. In this work, we propose an architecture and a sampling strategy to jointly model multi-view images and camera poses when fine-tuning a video generation model. Our core idea is a dual-stream architecture that attaches a dedicated pose generation model alongside a pre-trained video generation model via communication blocks, generating multi-view images and camera poses through separate streams. This design reduces interference between the pose and image modalities. Additionally, we propose an asynchronous sampling strategy that denoises camera poses faster than multi-view images, allowing rapidly denoised poses to condition multi-view generation, reducing mutual ambiguity and enhancing cross-modal consistency. Trained on multiple large-scale real-world datasets (RealEstate10K, MVImgNet, DL3DV-10K, ACID), VideoRFSplat outperforms existing text-to-3D direct generation methods that heavily depend on post-hoc refinement via score distillation sampling, achieving superior results without such refinement.

📄 PDF Abstract BibTeX arXiv:2503.15855

Code (0)

등록된 구현이 없습니다.

Tasks

3DGSText to 3DVideo Generation

Similar Papers 제목 키워드 기반

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

2025-12-29 · Tianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen 외 arxiv

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, wi…

Scene UnderstandingScene Generation

Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration

2026-06-29 · Jaewon Lee, Mangyu Kong, Euntai Kim arxiv

Merging multiple 3D Gaussian Splatting (3DGS) scenes into a single unified Gaussian representation is essential for large-scale 3D mapping and long-term map management. Despite its importance, this area remains underexpl…

LAYOUTDREAMER: Physics-guided Layout for Text-to-3D Compositional Scene Generation

2025-02-04 · Yang Zhou, Zongjin He, Qixuan Li, Chao Wang

Recently, the field of text-guided 3D scene generation has garnered significant attention. High-quality generation that aligns with physical realism and high controllability is crucial for practical 3D scene applications…

3DGSScene GenerationText to 3D

Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering

2023-11-30 · CVPR 2024 1 · Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli 외

Neural rendering methods have significantly advanced photo-realistic 3D scene rendering in various academic and industrial applications. The recent 3D Gaussian Splatting method has achieved the state-of-the-art rendering…

Neural Rendering

3D-USE: From Image-Level to Scene-Level Underwater Enhancement

2026-08-28 · Jieyu Yuan, Yuanlin Zhang, Jihong Li, Chunle Guo 외 arxiv

Underwater 3D reconstruction faithfully reproduces the color shifts and visibility loss of captured views, while physical inversion may leave estimation errors in the recovered scene appearance. We formulate Underwater S…

Image Enhancement3D Reconstruction