paper-with-me

홈 › Papers

MagiCapture: High-Resolution Multi-Concept Portrait Customization

2023-09-13 · Junha Hyung, Jaeyo Shin, Jaegul Choo

Large-scale text-to-image models including Stable Diffusion are capable of generating high-fidelity photorealistic portrait images. There is an active research area dedicated to personalizing these models, aiming to synthesize specific subjects or styles using provided sets of reference images. However, despite the plausible results from these personalization methods, they tend to produce images that often fall short of realism and are not yet on a commercially viable level. This is particularly noticeable in portrait image generation, where any unnatural artifact in human faces is easily discernible due to our inherent human bias. To address this, we introduce MagiCapture, a personalization method for integrating subject and style concepts to generate high-resolution portrait images using just a few subject and style references. For instance, given a handful of random selfies, our fine-tuned model can generate high-quality portrait images in specific styles, such as passport or profile photos. The main challenge with this task is the absence of ground truth for the composed concepts, leading to a reduction in the quality of the final output and an identity shift of the source subject. To address these issues, we present a novel Attention Refocusing loss coupled with auxiliary priors, both of which facilitate robust learning within this weakly supervised learning setting. Our pipeline also includes additional post-processing steps to ensure the creation of highly realistic outputs. MagiCapture outperforms other baselines in both quantitative and qualitative evaluations and can also be generalized to other non-human objects.

📄 PDF Abstract BibTeX arXiv:2309.06895

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationWeakly-supervised Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VToonify: Controllable High-Resolution Portrait Video Style Transfer

2022-09-22 · Shuai Yang, Liming Jiang, Ziwei Liu, Chen Change Loy

Generating high-quality artistic portrait videos is an important and desirable task in computer graphics and vision. Although a series of successful portrait image toonification models built upon the powerful StyleGAN ha…

Face AlignmentStyle TransferVideo Style TransferVocal Bursts Intensity Prediction

HeadsUp! High-Fidelity Portrait Image Super-Resolution

2025-10-10 · Renjie Li, Zihao Zhu, Xiaoyu Wang, Zhengzhong Tu arxiv

Portrait pictures, which typically feature both human subjects and natural backgrounds, are one of the most prevalent forms of photography on social media. Existing image super-resolution (ISR) techniques generally focus…

Image Super-Resolution

Pastiche Master: Exemplar-Based High-Resolution Portrait Style Transfer

2022-03-24 · CVPR 2022 1 · Shuai Yang, Liming Jiang, Ziwei Liu, Chen Change Loy

Recent studies on StyleGAN show high performance on artistic portrait generation by transfer learning with limited data. In this paper, we explore more challenging exemplar-based high-resolution portrait style transfer b…

Style TransferTransfer LearningVocal Bursts Intensity Prediction

Deep Single-Image Portrait Relighting

2019-10-01 · ICCV 2019 10 · Hao Zhou, Sunil Hadap, Kalyan Sunkavalli, David W. Jacobs

Conventional physically-based methods for relighting portrait images need to solve an inverse rendering problem, estimating face geometry, reflectance and lighting. However, the inaccurate estimation of face components c…

Inverse RenderingSingle-Image Portrait Relighting

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

2024-10-10 · Jiahao Cui, Hui Li, Yao Yao, Hao Zhu 외

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, int…

4kImage AnimationQuantizationVideo Generation