paper-with-me

Papers

MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling

2024-09-24 · CVPR 2025 1 · Yifang Men, Yuan YAO, Miaomiao Cui, Liefeng Bo

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

📄 PDF Abstract BibTeX arXiv:2409.16160

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

2026-04-19 · Junjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu 외 arxiv

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing a…

Visual StorytellingStory Continuation

Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

2023-11-28 · CVPR 2024 1 · Li Hu, Xin Gao, Peng Zhang, Ke Sun 외

Character Animation aims to generating character videos from still images through driving signals. Currently, diffusion models have become the mainstream in visual generation research, owing to their robust generative ca…

VideoComposer: Compositional Video Synthesis with Motion Controllability

2023-06-03 · NeurIPS 2023 11 · Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen 외

The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable video synthesis remains challenging due to t…

Image GenerationText-to-Video Generation

MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis

2024-10-28 · Di Qiu, Zheng Chen, Rui Wang, Mingyuan Fan 외

Recent advancements in character video synthesis still depend on extensive fine-tuning or complex 3D modeling processes, which can restrict accessibility and hinder real-time applicability. To address these challenges, w…

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

2026-04-10 · Haoyu Zhao, Zihao Zhang, Jiaxi Gu, Haoran Chen 외 arxiv

Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera control from text prompts or rely on labor…

Spatial ReasoningVideo Generation