paper-with-me

Papers

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

2026-03-16 · Minjun Kang, Inkyu Shin, Taeyeop Lee, Myungchul Kim, In So Kweon, Kuk-Jin Yoon arxiv

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.

📄 PDF Abstract BibTeX arXiv:2603.14965

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View Synthesis

Similar Papers 제목 키워드 기반

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

2026-04-01 · Jiyuan Hu, Zechuan Zhang, Zongxin Yang, Yi Yang arxiv

We present TRACE, a mesh-guided 3DGS editing framework that achieves automated, high-fidelity scene transformation. By anchoring video diffusion with explicit 3D geometry, TRACE uniquely enables fine-grained, part-level …

3D scene Editing

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

2026-03-17 · Tengjiao Yin, Jinglei Shi, Heng Guo, Xi Wang arxiv

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitati…

Reinforcement Learning

Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion

2026-01-12 · Farhad G. Zanjani, Hong Cai, Amirhossein Habibian arxiv

We present SetDiff, a geometry-grounded multi-view diffusion framework that enhances novel-view renderings produced by 3D Gaussian Splatting. Our method integrates explicit 3D priors, pixel-aligned coordinate maps and po…

Computational EfficiencyNovel View SynthesisAutonomous Driving

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

2025-04-01 · Tian-Xing Xu, Xiangjun Gao, WenBo Hu, Xiaoyu Li 외

Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstru…

4D reconstructionDepth Estimationparameter estimation

G$^2$TAM: Geometry Grounded Track Anything Model

2026-07-04 · Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu 외 arxiv

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explic…

Video Object SegmentationVideo SegmentationSpatial Reasoning3D Reconstruction