paper-with-me

홈 › Papers

FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization

2026-01-31 · Benxiang Zhai, Yifang Xu, Guofeng Zhang, Yang Li, Sidan Du arxiv

Benefiting from the significant advancements in text-to-image diffusion models, research in personalized image generation, particularly customized portrait generation, has also made great strides recently. However, existing methods either require time-consuming fine-tuning and lack generalizability or fail to achieve high fidelity in facial details. To address these issues, we propose FaceSnap, a novel method based on Stable Diffusion (SD) that requires only a single reference image and produces extremely consistent results in a single inference stage. This method is plug-and-play and can be easily extended to different SD models. Specifically, we design a new Facial Attribute Mixer that can extract comprehensive fused information from both low-level specific features and high-level abstract features, providing better guidance for image generation. We also introduce a Landmark Predictor that maintains reference identity across landmarks with different poses, providing diverse yet detailed spatial control conditions for image generation. Then we use an ID-preserving module to inject these into the UNet. Experimental results demonstrate that our approach performs remarkably in personalized and customized portrait generation, surpassing other state-of-the-art methods in this domain.

📄 PDF Abstract BibTeX arXiv:2602.00627

Code (0)

등록된 구현이 없습니다.

Tasks

Personalized Image Generation

Similar Papers 제목 키워드 기반

HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesis

2025-03-21 · CVPR 2025 1 · Mengtian Li, Jinshu Chen, Wanquan Feng, Bingchuan Li 외

Personalized portrait synthesis, essential in domains like social entertainment, has recently made significant progress. Person-wise fine-tuning based methods, such as LoRA and DreamBooth, can produce photorealistic outp…

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

2025-03-06 · Ziqi Ni, Ao Fu, Yi Zhou

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…

Audio-Visual Synchronization

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

2026-05-27 · He Feng, Yongjia Ma, Donglin Di, Lei Fan 외 arxiv

Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input granularity and motion accuracy. Existing…

Video Generation

SPF-Portrait: Towards Pure Portrait Customization with Semantic Pollution-Free Fine-tuning

2025-04-01 · Xiaole Xian, Zhichao Liao, Qingyu Li, Wenyu Qin 외

Fine-tuning a pre-trained Text-to-Image (T2I) model on a tailored portrait dataset is the mainstream method for text-driven customization of portrait attributes. Due to Semantic Pollution during fine-tuning, existing met…

Contrastive LearningIncremental Learning

IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models

2024-03-20 · Siying Cui, Jia Guo, Xiang An, Jiankang Deng 외

Leveraging Stable Diffusion for the generation of personalized portraits has emerged as a powerful and noteworthy tool, enabling users to create high-fidelity, custom character avatars based on their specific prompts. Ho…

DiversityImage GenerationPersonalized Image Generation