paper-with-me

Papers

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

2026-05-20 · Yunlong Wang, Jinjin Shi, Wenbin Gao, Xuran Xu, Runyu Shi, Ying Huang arxiv

Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthetics inherently inhibit one another. Supervised Fine-Tuning (SFT) is an effective method for enhancing the photorealism of image generation. However, it often leads to overfitting to the training dataset, corrupts pre-trained image priors, and degrades alignment or aesthetics. To break this bottleneck, we propose a feature supervision paradigm for Multimodal Diffusion Transformers (MM-DiT). Specifically, we introduce a lightweight cross-modal alignment mechanism that implicitly extracts multi-granularity vision-aligned text representations from SigLIP 2 and applies supervision to the image branch of MM-DiT during the training stage, with zero extra inference overhead. Our method injects vision-aligned text guidance while preserving the base model's original generalization, avoiding degradation caused by SFT. Furthermore, our method directly mines implicit multi-granularity aesthetic signals from pre-trained vision foundation models to optimize human-perceived aesthetics. Extensive experiments on MM-DiTs show that our method pushes the Pareto frontier and achieves synergistic improvements across text-image alignment, photorealism, and human-perceived aesthetics.

📄 PDF Abstract BibTeX arXiv:2605.20640

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait

2025-01-28 · Han Yang, Enis Simsar, Sotiris Anagnostidis, Yanlong Zang 외

Existing diffusion models show great potential for identity-preserving generation. However, personalized portrait generation remains challenging due to the diversity in user profiles, including variations in appearance a…

Representation Learning

FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation

2026-02-25 · Weiting Tan, Andy T. Liu, Ming Tu, Xinghua Qu 외 arxiv

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propos…

Reinforcement LearningVideo Generation

Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior

2024-04-16 · Yiqian Wu, Hao Xu, Xiangjun Tang, Xien Chen 외

Existing neural rendering-based text-to-3D-portrait generation methods typically make use of human geometry prior and diffusion models to obtain guidance. However, relying solely on geometry information introduces issues…

Neural RenderingText to 3D

HeadsUp! High-Fidelity Portrait Image Super-Resolution

2025-10-10 · Renjie Li, Zihao Zhu, Xiaoyu Wang, Zhengzhong Tu arxiv

Portrait pictures, which typically feature both human subjects and natural backgrounds, are one of the most prevalent forms of photography on social media. Existing image super-resolution (ISR) techniques generally focus…

Image Super-Resolution

PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

2026-04-04 · Yuyang Sha, Zijie Lou, Youyun Tang, Xiaochao Qu 외 arxiv

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrai…

Visual Question Answering