paper-with-me

Papers

Reangle-A-Video: 4D Video Generation as Video-to-Video Translation

2025-03-12 · Hyeonho Jeong, Suhyeon Lee, Jong Chul Ye

We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the multi-view video generation task as video-to-videos translation, leveraging publicly available image and video diffusion priors. In essence, Reangle-A-Video operates in two stages. (1) Multi-View Motion Learning: An image-to-video diffusion transformer is synchronously fine-tuned in a self-supervised manner to distill view-invariant motion from a set of warped videos. (2) Multi-View Consistent Image-to-Images Translation: The first frame of the input video is warped and inpainted into various camera perspectives under an inference-time cross-view consistency guidance using DUSt3R, generating multi-view consistent starting images. Extensive experiments on static view transport and dynamic camera control show that Reangle-A-Video surpasses existing methods, establishing a new solution for multi-view video generation. We will publicly release our code and data. Project page: https://hyeonho99.github.io/reangle-a-video/

📄 PDF Abstract BibTeX arXiv:2503.09151

Code (0)

등록된 구현이 없습니다.

Tasks

TranslationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning

2024-11-07 · CVPR 2025 1 · David Junhao Zhang, Roni Paiss, Shiran Zada, Nikhil Karnad 외

Recently, breakthroughs in video modeling have allowed for controllable camera trajectories in generated videos. However, these methods cannot be directly applied to user-provided videos that are not generated by a video…

Multi-sentence Video Grounding for Long Video Generation

2024-07-18 · Wei Feng, Xin Wang, Hong Chen, Zeyang Zhang 외

Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal consistency of generated videos and the h…

Moment RetrievalRetrievalSentenceVideo Editing+2

LVD-2M: A Long-take Video Dataset with Temporally Dense Captions

2024-10-14 · Tianwei Xiong, Yuqing Wang, Daquan Zhou, Zhijie Lin 외

The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest…

Video CaptioningVideo Generation

LongCat-Video Technical Report

2025-10-25 · Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang 외 arxiv

Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameter…

Video Generation

Video Generation Beyond a Single Clip

2023-04-15 · Hsin-Ping Huang, Yu-Chuan Su, Ming-Hsuan Yang

We tackle the long video generation problem, i.e.~generating videos beyond the output length of video generation models. Due to the computation resource constraints, video generation models can only generate video clips …

Video Generation