paper-with-me

홈 › Papers

MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

2024-08-27 · Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, Haoqiang Fan, Jin Wang

Diffusion models have demonstrated superior performance in the field of portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality. To address this issue, we introduce MegActor-$\Sigma$: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation. Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework. To further achieve flexible combinations of mixed-modal control signals, we propose a `Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the `Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality. Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset to train MegActor-$\Sigma$. Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations, outperforming previous methods trained on private dataset.

📄 PDF Abstract BibTeX arXiv:2408.14975

Code (2)

megvii-research/megactor pytorch
megvii-research/megfaceanimate pytorch

Tasks

Portrait Animation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows

2025-10-03 · John Nguyen, Marton Havasi, Tariq Berrada, Luke Zettlemoyer 외 arxiv

We present OneFlow, the first non-autoregressive multimodal model that enables variable-length and concurrent mixed-modal generation. Unlike autoregressive models that enforce rigid causal ordering between text and image…

Image Generation

MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation

2025-02-09 · Zhifei Yang, Keyang Lu, Chao Zhang, Jiaxing Qi 외

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs …

Scene Generation

OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control

2026-04-07 · Yukun Wang, Ruihuang Li, Jiale Tao, Shiyuan Yang 외 arxiv

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent …

Multi-Task LearningVideo Generation

Perimeter control in a mixed bimodal bathtub model

2022-08-09 · Takao Dantsuji, Yuki Takayama, Daisuke Fukuda

Perimeter control involves monitoring network-wide traffic and regulating traffic inflow to alleviate hypercongestion. Implementation of transit priority with perimeter control measures, which allow transit into a contro…

model

Unified Force and Motion Adaptive-Integral Control of Flexible Robot Manipulators

2023-09-18 · Carlos R. de Cos, José Ángel Acosta

In this paper, an adaptive nonlinear strategy for the motion and force control of flexible manipulators is proposed. The approach provides robust motion control until contact is detected when force control is then availa…

Position