paper-with-me

Papers

Phantom: Subject-consistent video generation via cross-modal alignment

2025-02-16 · Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Qian He, Xinglong Wu

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent video through textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages. The project homepage is here https://phantom-video.github.io/Phantom/.

📄 PDF Abstract BibTeX arXiv:2502.11079

Code (1)

phantom-video/phantom pytorch

Tasks

cross-modal alignmentHuman-Domain Subject-to-VideoOpen-Domain Subject-to-VideoSingle-Domain Subject-to-VideoTripletVideo Generation

Similar Papers 제목 키워드 기반

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

2025-10-16 · Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao 외 arxiv

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent ident…

Reinforcement LearningVideo Generation

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

2026-04-09 · Ying Shen, Jerry Xiong, Tianjiao Yu, Ismini Lourentzou arxiv

Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However, emerging evidence suggests that simply scaling data and model size…

Video Generation

Cell Phantom Video Generation in Elliptical Fourier Descriptor Domain

2026-05-21 · Francesco Benedetto, Roberto Basla, Luca Magri, Giacomo Boracchi arxiv

Training Deep Neural Networks for tracking individual cells in biomedical videos requires a large amount of annotated data. The annotation of videos for cell tracking is very time consuming and often requires domain expe…

Video Generation

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

2025-10-01 · Zhaoyang Li, Dongjun Qian, Kai Su, Qishuai Diao 외 arxiv

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall …

Video Generation

MV-S2V: Multi-View Subject-Consistent Video Generation

2026-01-25 · Ziyang Song, Xinyu Gong, Bangya Liu, Zelin Zhao arxiv

Existing Subject-to-Video Generation (S2V) methods have achieved high-fidelity and subject-consistent video generation, yet remain constrained to single-view subject references. This limitation renders the S2V task reduc…

Video Generation