paper-with-me

Papers

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

2026-04-21 · Chaonan Ji, Jinwei Qi, Sheng Xu, Peng Zhang, Bang Zhang arxiv

Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods designed for control may struggle with fidelity and robust disentanglement. Instead of treating facial motion as a monolithic signal, we explore an alternative compositional perspective. In this paper, we introduce PortraitDirector, a novel framework that formulates face reenactment as a hierarchical composition task, achieving high-fidelity and controllable results. We employ a Hierarchical Motion Disentanglement and Composition strategy, deconstructing facial motion into a Spatial Layer for physical movements and a Semantic Layer for emotional content. The Spatial Layer comprises: (i) global head pose, managed via a dedicated representation and injection pathway; (ii) spatially separated local facial expressions, distilled from cropped facial regions and purged of emotional cues via Emotion-Filtering Module leveraging an information bottleneck. The Semantic Layer contains a derived global emotion. The disentangled components are then recomposed into an expressive motion latent. Furthermore, we engineer the framework for real-time performance through a suite of optimizations, including diffusion distillation, causal attention and VAE acceleration. PortraitDirector achieves streaming, high-fidelity, controllable 512 x 512 face reenactment at 20 FPS with a end-to-end 800 ms latency on a single 5090 GPU.

📄 PDF Abstract BibTeX arXiv:2604.19129

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Image-to-image Translation via Hierarchical Style Disentanglement

2021-03-02 · CVPR 2021 1 · Xinyang Li, Shengchuan Zhang, Jie Hu, Liujuan Cao 외

Recently, image-to-image translation has made significant progress in achieving both multi-label (\ie, translation conditioned on different labels) and multi-style (\ie, generation with diverse styles) tasks. However, du…

DisentanglementImage-to-Image TranslationMultimodal Unsupervised Image-To-Image TranslationTranslation

Controllable Data Generation Via Iterative Data-Property Mutual Mappings

2023-10-11 · Bo Pan, Muran Qin, Shiyu Wang, Yifei Zhang 외

Deep generative models have been widely used for their ability to generate realistic data samples in various areas, such as images, molecules, text, and speech. One major goal of data generation is controllability, namel…

Disentanglement

HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation

2024-10-18 · Bo Cheng, Yuhang Ma, Liebucha Wu, Shanyuan Liu 외

The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases inclu…

DisentanglementImage GenerationLayout GenerationLayout-to-Image Generation

Semi-Supervised StyleGAN for Disentanglement Learning

2020-03-06 · ICML 2020 1 · Weili Nie, Tero Karras, Animesh Garg, Shoubhik Debnath 외

Disentanglement learning is crucial for obtaining disentangled representations and controllable generation. Current disentanglement methods face several inherent limitations: difficulty with high-resolution images, prima…

DisentanglementRepresentation Learning

GCVAE: Generalized-Controllable Variational AutoEncoder

2022-06-09 · Kenneth Ezukwoke, Anis Hoayek, Mireille Batton-Hubert, Xavier Boucher

Variational autoencoders (VAEs) have recently been used for unsupervised disentanglement learning of complex density distributions. Numerous variants exist to encourage disentanglement in latent space while improving rec…

Disentanglement