paper-with-me

Papers

MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation

2024-11-07 · Han Yang, Sotiris Anagnostidis, Enis Simsar, Thomas Hofmann

We propose MegaPortrait. It's an innovative system for creating personalized portrait images in computer vision. It has three modules: Identity Net, Shading Net, and Harmonization Net. Identity Net generates learned identity using a customized model fine-tuned with source images. Shading Net re-renders portraits using extracted representations. Harmonization Net fuses pasted faces and the reference image's body for coherent results. Our approach with off-the-shelf Controlnets is better than state-of-the-art AI portrait products in identity preservation and image fidelity. MegaPortrait has a simple but effective design and we compare it with other methods and products to show its superiority.

📄 PDF Abstract BibTeX arXiv:2411.04357

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion

2026-01-29 · Pu Cao, Yiyang Ma, Feng Zhou, Xuedan Yin 외 arxiv

In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID). In recent ImageNet-scale AE studies, we…

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

2026-02-26 · Tianxing Xu, Zixuan Wang, Guangyuan Wang, Li Hu 외 arxiv

World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficulties in two key areas: maintaining long-term content consistency when sce…

3D ReconstructionVideo Generation

NFCDS: A Plug-and-Play Noise Frequency-Controlled Diffusion Sampling Strategy for Image Restoration

2026-01-29 · Zhen Wang, Hongyi Liu, Jianing Li, Zhihui Wei arxiv

Diffusion sampling-based Plug-and-Play (PnP) methods produce images with high perceptual quality but often suffer from reduced data fidelity, primarily due to the noise introduced during reverse diffusion. To address thi…

Image Restoration

Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution

2026-03-14 · Dan Wang, Haiyan Sun, Shan Du, Z. Jane Wang 외 arxiv

Image super-resolution (SR) aims to reconstruct high resolution images with both high perceptual quality and low distortion, but is fundamentally limited by the perception-distortion trade-off. GAN-based SR methods reduc…

Image Super-Resolution

Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis

2023-04-07 · ICCV 2023 1 · Qiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui 외

Diffusion-based models have achieved state-of-the-art performance on text-to-image synthesis tasks. However, one critical limitation of these models is the low fidelity of generated images with respect to the text descri…

DenoisingImage Generation