paper-with-me

Papers

View-Consistent Diffusion Representations for 3D-Consistent Video Generation

2025-11-24 · Duolikun Danier, Ge Gao, Steven McDonagh, Changjian Li, Hakan Bilen, Oisin Mac Aodha arxiv

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D inconsistencies, e.g., objects and structures deforming under changes in camera pose, which can undermine user experience and simulation fidelity. Motivated by recent findings on representation alignment for diffusion models, we hypothesize that improving the multi-view consistency of video diffusion representations will yield more 3D-consistent video generation. Through detailed analysis on multiple recent camera-controlled video diffusion models we reveal strong correlations between 3D-consistent representations and videos. We also propose ViCoDR, a new approach for improving the 3D consistency of video models by learning multi-view consistent diffusion representations. We evaluate ViCoDR on camera controlled image-to-video, text-to-video, and multi-view generation models, demonstrating significant improvements in the 3D consistency of the generated videos. Project page: https://danier97.github.io/ViCoDR.

📄 PDF Abstract BibTeX arXiv:2511.18991

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

2025-04-03 · CVPR 2025 1 · Shengjun Zhang, Jinzhao Li, Xin Fei, Hao liu 외

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suff…

Scene GenerationVideo Generation

Human-VDM: Learning Single-Image 3D Human Gaussian Splatting from Video Diffusion Models

2024-09-04 · Zhibin Liu, Haoye Dong, Aviral Chharia, Hefeng Wu

Generating lifelike 3D humans from a single RGB image remains a challenging task in computer vision, as it requires accurate modeling of geometry, high-quality texture, and plausible unseen parts. Existing methods typica…

Lifelike 3D Human Generation

ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models

2023-12-03 · CVPR 2024 1 · Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko 외

Generating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and rendering high-quality, spatially consistent new …

Novel View SynthesisObject

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

2026-08-24 · Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan arxiv

Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target views lie far from th…

Point Clouds

GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation

2025-03-08 · Ye Tao, Jiawei Zhang, Yahao Shi, Dongqing Zou 외

Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models a…

3D GenerationDecoderVideo Generation