paper-with-me

홈 › Papers

Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

2024-05-21 · Yue Han, Junwei Zhu, Keke He, Xu Chen, Yanhao Ge, Wei Li, Xiangtai Li, Jiangning Zhang, Chengjie Wang, Yong liu

Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and the results have not yet achieved satisfactory performance levels. To address this issue, we introduce Face-Adapter, an efficient and effective adapter designed for high-precision and high-fidelity face editing for pre-trained diffusion models. We observe that both face reenactment/swapping tasks essentially involve combinations of target structure, ID and attribute. We aim to sufficiently decouple the control of these factors to achieve both tasks in one model. Specifically, our method contains: 1) A Spatial Condition Generator that provides precise landmarks and background; 2) A Plug-and-play Identity Encoder that transfers face embeddings to the text space by a transformer decoder. 3) An Attribute Controller that integrates spatial conditions and detailed attributes. Face-Adapter achieves comparable or even superior performance in terms of motion control precision, ID retention capability, and generation quality compared to fully fine-tuned face reenactment/swapping models. Additionally, Face-Adapter seamlessly integrates with various StableDiffusion models.

📄 PDF Abstract BibTeX arXiv:2405.12970

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDecoderFace Reenactment

Methods 이 논문이 사용한 방법론

Adapter 설명 없음
Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

2024-07-23 · Fang-Duo Tsai, Shih-Lun Wu, Haven Kim, Bo-Yu Chen 외

Text-to-music models allow users to generate nearly realistic musical audio with textual commands. However, editing music audios remains challenging due to the conflicting desiderata of performing fine-grained alteration…

AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation

2026-05-17 · Yixuan Han arxiv

We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editing conditions. Instead of relying on large-scale vision-language mod…

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

2026-05-28 · Hao Wu, Xiangyang Luo, Hao Wang, Jiawei Zhang 외 arxiv

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets…

Talking Face Generation

ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion

2025-10-06 · Foivos Paraperas Papantoniou, Stefanos Zafeiriou arxiv

Human-centric generative models designed for AI-driven storytelling must bring together two core capabilities: identity consistency and precise control over human performance. While recent diffusion-based approaches have…

Lynx: Towards High-Fidelity Personalized Video Generation

2025-09-19 · Shen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu 외 arxiv

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure …

Video Generation