paper-with-me

Papers

OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control

2026-04-07 · Yukun Wang, Ruihuang Li, Jiale Tao, Shiyuan Yang, Liyi Chen, Zhantao Yang, Handz, Yulan Guo, Shuai Shao, Qinglin Lu arxiv

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent control. In this work, we introduce OmniCamera, a unified framework designed to explicitly disentangle and command these two dimensions. This compositional approach enables flexible video generation by allowing arbitrary pairings of camera and content conditions, unlocking unprecedented creative control. To overcome the fundamental challenges of modality conflict and data scarcity inherent in such a system, we present two key innovations. First, we construct OmniCAM, a novel hybrid dataset combining curated real-world videos with synthetic data that provides diverse paired examples for robust multi-task learning. Second, we propose a Dual-level Curriculum Co-Training strategy that mitigates modality interference and synergistically learns from diverse data sources. This strategy operates on two levels: first, it progressively introduces control modalities by difficulties (condition-level), and second, trains for precise control on synthetic data before adapting to real data for photorealism (data-level). As a result, OmniCamera achieves state-of-the-art performance, enabling flexible control for complex camera movements while maintaining superior visual quality.

📄 PDF Abstract BibTeX arXiv:2604.06010

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningVideo Generation

Similar Papers 제목 키워드 기반

Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

2024-09-27 · Min Yang, Zichen Zhang, LiMin Wang

With the development of video understanding, there is a proliferation of tasks for clip-level temporal video analysis, including temporal action detection (TAD), temporal action segmentation (TAS), and generic event boun…

Action DetectionAction SegmentationBoundary DetectionGeneric Event Boundary Detection+3

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

2026-02-10 · Jialun Liu, Tian Li, Xiao Cao, Yukuo Ma 외 arxiv

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, l…

Text-to-Video Generation

UniVBench: Towards Unified Evaluation for Video Foundation Models

2026-02-25 · Jianhui Wei, Xiaotian Zhang, Yichen Li, Yuan Wang 외 arxiv

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, exi…

Instruction FollowingVideo ReconstructionVideo Generation

LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling

2022-06-14 · CVPR 2023 1 · Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin 외

Unified vision-language frameworks have greatly advanced in recent years, most of which adopt an encoder-decoder architecture to unify image-text tasks as sequence-to-sequence generation. However, existing video-language…

DecoderLanguage ModelingLanguage ModellingMasked Language Modeling+6

A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking

2025-05-26 · Zixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui 외

The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Uni…

BenchmarkingOptical Flow EstimationSynthetic Data Generation