paper-with-me

Papers

CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

2024-11-27 · CVPR 2025 1 · Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, Aleksander Holynski

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera poses and timestamps. Combined with a novel sampling approach, this model can transform a single monocular video into a multi-view video, enabling robust 4D reconstruction via optimization of a deformable 3D Gaussian representation. We demonstrate competitive performance on novel view synthesis and dynamic scene reconstruction benchmarks, and highlight the creative capabilities for 4D scene generation from real or generated videos. See our project page for results and interactive demos: https://cat-4d.github.io/.

📄 PDF Abstract BibTeX arXiv:2411.18613

Code (0)

등록된 구현이 없습니다.

Tasks

4D reconstructionNovel View SynthesisScene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CAT3D: Create Anything in 3D with Multi-View Diffusion Models

2024-05-16 · Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee 외

Advances in 3D reconstruction have enabled high-quality 3D capture, but require a user to collect hundreds to thousands of images to create a 3D scene. We present CAT3D, a method for creating anything in 3D by simulating…

3D Reconstruction

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

2026-08-24 · Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan arxiv

Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views are sparse or target views lie far from th…

Point Clouds

Image Coding for Machines with Edge Information Learning Using Segment Anything

2024-03-07 · Takahiro Shindo, Kein Yamada, Taiju Watanabe, Hiroshi Watanabe

Image Coding for Machines (ICM) is an image compression technique for image recognition. This technique is essential due to the growing demand for image recognition AI. In this paper, we propose a method for ICM that foc…

Image CompressionVideo Compression

SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion

2025-02-17 · Junxian Ma, Shiwen Wang, Jian Yang, Junyi Hu 외

Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization. However, existing methods typically rely on constrained audio-visual alignment priors or multi-stage learning of i…

Motion Synthesis

EraseAnything++: Enabling Concept Erasure in Rectified Flow Transformers Leveraging Multi-Object Optimization

2026-03-01 · Zhaoxin Fan, Nanxiang Jiang, Daiheng Gao, Shiji Zhou 외 arxiv

Removing undesired concepts from large-scale text-to-image (T2I) and text-to-video (T2V) diffusion models while preserving overall generative quality remains a major challenge, particularly as modern models such as Stabl…

Video Generation