paper-with-me

홈 › Papers

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

2025-12-26 · Run Ling, Ke Cao, Jian Lu, Ao Ma, Haowei Liu, Runze He, Changwei Wang, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Peng Wu, Guibing Guo, Wei Feng, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Xingwei Wang arxiv

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.

📄 PDF Abstract BibTeX arXiv:2512.22310

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

2022-12-08 · CVPR 2023 1 · Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, Christian Theobalt

Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality. In response to these limitations, we introduce MoFusion, i.e., a new de…

DenoisingDiversityMotion Synthesis

StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework

2024-05-09 · Yiheng Huang, Hui Yang, Chuanchen Luo, Yuxi Wang 외

Thanks to the powerful generative capacity of diffusion models, recent years have witnessed rapid progress in human motion generation. Existing diffusion-based methods employ disparate network architectures and training …

DenoisingMotion Generation

DemoFusion: Democratising High-Resolution Image Generation With No $$$

2023-11-24 · CVPR 2024 1 · Ruoyi Du, Dongliang Chang, Timothy Hospedales, Yi-Zhe Song 외

High-resolution image generation with Generative Artificial Intelligence (GenAI) has immense potential but, due to the enormous capital investment required for training, it is increasingly centralised to a few large corp…

Image Generation

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

2024-04-11 · ICCV 2023 1 · Shan Wang, Chuong Nguyen, Jiawei Liu, Kaihao Zhang 외

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, s…

Autonomous DrivingSegmentation

Pretrained Diffusion Models for Unified Human Motion Synthesis

2022-12-06 · Jianxin Ma, Shuai Bai, Chang Zhou

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a m…

Motion GenerationMotion SynthesisOpen-Ended Question Answering