paper-with-me

Papers

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

2025-04-15 · Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang, Chi Zhang, Xuelong Li

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff, aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. This allows flexible manipulation of each modality's role, enabling support for a wide range of tasks. Consequently, our model supports three key functionalities: (1) Text-conditioned video generation: multi-modal visual video sequences (i.e., rgb, depth, canny, segmentaion) are generated based on the text conditions in one diffusion process; (2) Video understanding: OmniVDiff can estimate the depth, canny map, and semantic segmentation across the input rgb frames while ensuring coherence with the rgb input; and (3) X-conditioned video generation: OmniVDiff generates videos conditioned on fine-grained attributes (e.g., depth maps or segmentation maps). By integrating these diverse tasks into a unified video diffusion framework, OmniVDiff enhances the flexibility and scalability for controllable video diffusion, making it an effective tool for a variety of downstream applications, such as video-to-video translation. Extensive experiments demonstrate the effectiveness of our approach, highlighting its potential for various video-related applications.

📄 PDF Abstract BibTeX arXiv:2504.10825

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationVideo GenerationVideo Understanding

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

2026-05-19 · Hongji Yang, Songlian Li, Yucheng Zhou, Xiaotong Zhao 외 arxiv

Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in professional production workflows such …

Reinforcement LearningVideo Generation

OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

2024-12-12 · Weiqi Li, Shijie Zhao, Chong Mou, Xuhan Sheng 외

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, th…

Image to Video GenerationVideo Generation

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

2025-06-02 · Sen Liang, Zhentao Yu, Zhengguang Zhou, Teng Hu 외

The emergence of Diffusion Transformers (DiT) has brought significant advancements to video generation, especially in text-to-video and image-to-video tasks. Although video generation is widely applied in various fields,…

Data AugmentationHuman AnimationVideo Generation

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

2026-02-12 · Xu Guo, Fulong Ye, Qichao Sun, Liyang Chen 외 arxiv

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video e…

Video Generation

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

2026-03-09 · Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee 외 arxiv

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlik…

Cross-Modal RetrievalSpeech RecognitionImage Generation