paper-with-me

Papers

Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing

2026-02-09 · Hao Yang, Zhiyu Tan, Jia Gong, Luozheng Qin, Hesen Chen, Xiaomeng Yang, Yuqing Sun, Yuetan Lin, Mengping Yang, Hao Li arxiv

We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified video generation and editing. Our key idea is to exploit the understanding and reasoning capabilities of MLLMs to produce explicit target captions to interpret user instructions. In this way, the rich contextual representations from the understanding model are directly used to guide the generative process, thereby improving performance on complex and compositional editing. Moreover, a lightweight adapter is developed to inject multimodal conditional tokens into pretrained text-to-video diffusion models, allowing maximum reuse of their powerful generative priors in a parameter-efficient manner. Benefiting from these designs, we scale up Omni-Video 2 to a 14B video diffusion model on meticulously curated training data with quality, supporting high quality text-to-video generation and various video editing tasks such as object removal, addition, background change, complex motion editing, \emph{etc.} We evaluate the performance of Omni-Video 2 on the FiVE benchmark for fine-grained video editing and the VBench benchmark for text-to-video generation. The results demonstrate its superior ability to follow complex compositional instructions in video editing, while also achieving competitive or superior quality in video generation tasks.

📄 PDF Abstract BibTeX arXiv:2602.08820

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Omni-Video: Democratizing Unified Video Understanding and Generation

2025-07-08 · Zhiyu Tan, Hao Yang, Luozheng Qin, Jia Gong 외

Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on…

Video GenerationVideo Understanding

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

2025-04-15 · Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 외

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff, aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats al…

Semantic SegmentationVideo GenerationVideo Understanding

OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis

2025-12-11 · Xiang Fan, Sharath Girish, Vivek Ramanujan, Chaoyang Wang 외 arxiv

Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore…

Novel View SynthesisVideo Generation

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

2025-02-03 · Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng 외

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generatio…

Human AnimationHuman-Object Interaction DetectionMotion GenerationVideo Generation

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

2025-10-12 · Caorui Li, Yu Chen, Yiyan Ji, Jin Xu 외 arxiv

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities…

Causal InferenceVisual Reasoning