paper-with-me

홈 › Papers

UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution

2025-10-09 · Shian Du, Menghan Xia, Chang Liu, Quande Liu, Xintao Wang, Pengfei Wan, Xiangyang Ji arxiv

Cascaded video super-resolution has emerged as a promising technique for decoupling the computational burden associated with generating high-resolution videos using large foundation models. Existing studies, however, are largely confined to text-to-video tasks and fail to leverage additional generative conditions beyond text, which are crucial for ensuring fidelity in multi-modal video generation. We address this limitation by presenting UniMMVSR, the first unified generative video super-resolution framework to incorporate hybrid-modal conditions, including text, images, and videos. We conduct a comprehensive exploration of condition injection strategies, training schemes, and data mixture techniques within a latent video diffusion model. A key challenge was designing distinct data construction and condition utilization methods to enable the model to precisely utilize all condition types, given their varied correlations with the target video. Our experiments demonstrate that UniMMVSR significantly outperforms existing methods, producing videos with superior detail and a higher degree of conformity to multi-modal conditions. We also validate the feasibility of combining UniMMVSR with a base model to achieve multi-modal guided generation of 4K video, a feat previously unattainable with existing techniques.

📄 PDF Abstract BibTeX arXiv:2510.08143

Code (0)

등록된 구현이 없습니다.

Tasks

Video Super-ResolutionVideo Generation

Similar Papers 제목 키워드 기반

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

2025-06-25 · Yanzhe Chen, Huasong Zhong, Yan Li, Zhenheng Yang

Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existin…

16k

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

2026-07-26 · Zhijing Cheng, Xuancheng Zhang, Donglin Di, Lei Fan 외 arxiv

End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories cond…

Instruction FollowingAutonomous Driving

Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection

2026-05-28 · Yangchen Wu, Huiqiang Xie arxiv

Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simul…

Multi-class Anomaly Detection

Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion

2025-12-15 · Toan Le Ngo Thanh, Phat Ha Huu, Tan Nguyen Dang Duy, Thong Nguyen Le Minh 외 arxiv

The exponential growth of video content has created an urgent need for efficient multimodal moment retrieval systems. However, existing approaches face three critical challenges: (1) fixed-weight fusion strategies fail a…

Moment Retrieval

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

2026-04-26 · Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin 외 arxiv

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attenti…

Video Generation