paper-with-me

Papers

MV-S2V: Multi-View Subject-Consistent Video Generation

2026-01-25 · Ziyang Song, Xinyu Gong, Bangya Liu, Zelin Zhao arxiv

Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to single-view subject references. This limitation renders the S2V task reducible to an S2I + I2V pipeline, failing to exploit the full potential of video subject control. In this work, we propose and address the challenging Multi-View S2V (MV-S2V) task, which synthesizes videos from multiple reference views to enforce 3D-level subject consistency. Regarding the scarcity of training data, we first develop a synthetic data curation pipeline to generate highly customized synthetic data, complemented by a small-scale real-world captured dataset to boost the training of MV-S2V. Another key issue lies in the potential confusion between cross-subject and cross-view references in conditional generation. To overcome this, we further introduce Temporally Shifted RoPE (TS-RoPE) to distinguish between different subjects and distinct views of the same subject in reference conditioning. Our framework achieves superior 3D subject consistency w.r.t. multi-view reference images and high-quality visual outputs, establishing a new meaningful direction for subject-driven video generation. Code and data are available at: https://szy-young.github.io/mv-s2v

📄 PDF Abstract BibTeX arXiv:2601.17756

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

HumanOrbit: 3D Human Reconstruction as 360° Orbit Generation

2026-02-27 · Keito Suzuki, Kunyao Chen, Lei Wang, Bang Du 외 arxiv

We present a method for generating a full 360° orbit video around a person from a single input image. Existing methods typically adapt image-based diffusion models for multi-view synthesis, but yield inconsistent results…

3D Human ReconstructionImage Generation

Phantom: Subject-consistent video generation via cross-modal alignment

2025-02-16 · Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen 외

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video,…

cross-modal alignmentHuman-Domain Subject-to-VideoOpen-Domain Subject-to-VideoSingle-Domain Subject-to-Video+2

3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model

2026-03-19 · Hyun-kyu Ko, Jihyeon Park, Younghyun Kim, Dongheok Park 외 arxiv

Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and next-generation e-commerce. However, de…

Video Generation

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

2025-10-01 · Zhaoyang Li, Dongjun Qian, Kai Su, Qishuai Diao 외 arxiv

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall …

Video Generation

Memento: Reconstruct to Remember for Consistent Long Video Generation

2026-06-12 · Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu 외 arxiv

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating vide…

Video Generation