paper-with-me

홈 › Papers

VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction

2026-05-14 · Kaixin Zhu, Yiwen Tang, Yifan Yang, Renrui Zhang, Bohan Zeng, Ziyu Guo, Ruichuan An, Zhou Liu, Qizhi Chen, Delin Qu, Jaehong Yoon, Wentao Zhang arxiv

High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environments in a single forward pass. However, despite their strong performance in static scene perception, these models remain limited in responding to dynamic human instructions, which restricts their use in interactive applications. Existing editing methods typically rely on a 2D-lifting strategy, where individual views are edited independently and then lifted back into 3D space. This indirect pipeline often leads to blurry textures and inconsistent geometry, as 2D editors lack the spatial awareness required to preserve structure across viewpoints. To address these limitations, we propose VGGT-Edit, a feed-forward framework for text-conditioned native 3D scene editing. VGGT-Edit introduces depth-synchronized text injection to align semantic guidance with the backbone's spatial poses, ensuring stable instruction grounding. This semantic signal is then processed by a residual transformation head, which directly predicts 3D geometric displacements to deform the scene while preserving background stability. To ensure high-fidelity results, we supervise the framework with a multi-term objective function that enforces geometric accuracy and cross-view consistency. We also construct the DeltaScene Dataset, a large-scale dataset generated through an automated pipeline with 3D agreement filtering to ensure ground-truth quality. Experiments show that VGGT-Edit substantially outperforms 2D-lifting baselines, producing sharper object details, stronger multi-view consistency, and near-instant inference speed. The project page is https://chriszkxxx.github.io/VGGT-Edit/.

📄 PDF Abstract BibTeX arXiv:2605.15186

Code (0)

등록된 구현이 없습니다.

Tasks

3D scene Editing

Similar Papers 제목 키워드 기반

VGGT: Visual Geometry Grounded Transformer

2025-03-14 · CVPR 2025 1 · Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi 외

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. T…

Depth EstimationNovel View Synthesisparameter estimationPoint cloud reconstruction+1

DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving

2026-03-09 · Zhuolin He, Jing Li, Guanghao Li, Xiaolei Chen 외 arxiv

Dynamic scene reconstruction in autonomous driving remains a fundamental challenge due to significant temporal variations, moving objects, and complex scene dynamics. Existing feed-forward 3D models have demonstrated str…

Autonomous Driving

VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction

2026-01-27 · Dominic Maggio, Luca Carlone arxiv

We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree-of-freedo…

Object DetectionImage Retrieval

FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention

2025-12-01 · Zipeng Wang, Dan Xu arxiv

3D reconstruction from multi-view images is a core challenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to traditional per-scene optimization techniques. Among th…

3D Reconstruction

PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception

2025-10-20 · Kaichen Zhou, Yuhan Wang, Grace Chen, Xinhai Chang 외 arxiv

Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datase…

Camera Pose EstimationDepth Estimation