paper-with-me

홈 › Papers

ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation

2025-07-16 · Hyun-Jun Jin, Young-Eun Kim, Seong-Whan Lee arxiv

Recently, personalized portrait generation with a text-to-image diffusion model has significantly advanced with Textual Inversion, emerging as a promising approach for creating high-fidelity personalized images. Despite its potential, current Textual Inversion methods struggle to maintain consistent facial identity due to semantic misalignments between textual and visual embedding spaces regarding identity. We introduce ID-EA, a novel framework that guides text embeddings to align with visual identity embeddings, thereby improving identity preservation in a personalized generation. ID-EA comprises two key components: the ID-driven Enhancer (ID-Enhancer) and the ID-conditioned Adapter (ID-Adapter). First, the ID-Enhancer integrates identity embeddings with a textual ID anchor, refining visual identity embeddings derived from a face recognition model using representative text embeddings. Then, the ID-Adapter leverages the identity-enhanced embedding to adapt the text condition, ensuring identity preservation by adjusting the cross-attention module in the pre-trained UNet model. This process encourages the text features to find the most related visual clues across the foreground snippets. Extensive quantitative and qualitative evaluations demonstrate that ID-EA substantially outperforms state-of-the-art methods in identity preservation metrics while achieving remarkable computational efficiency, generating personalized portraits approximately 15 times faster than existing approaches.

📄 PDF Abstract BibTeX arXiv:2507.11990

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationComputational EfficiencyFace Recognition

Similar Papers 제목 키워드 기반

Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation

2026-07-08 · Wenyan Xu, Alizer Wong arxiv

Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and ma…

Text-to-Image Generation

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

2025-06-09 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang 외

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyViv…

Video Generation

Contextual Categorization Enhancement through LLMs Latent-Space

2024-04-25 · Zineddine Bettouche, Anas Safi, Andreas Fischer

Managing the semantic quality of the categorization in large textual datasets, such as Wikipedia, presents significant challenges in terms of complexity and cost. In this paper, we propose leveraging transformer models t…

Dimensionality Reduction

TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement

2025-09-08 · Jibai Lin, Bo Ma, Yating Yang, Xi Zhou 외 arxiv

Subject-driven image generation (SDIG) aims to manipulate specific subjects within images while adhering to textual instructions, a task crucial for advancing text-to-image diffusion models. SDIG requires reconciling the…

Image Generation

Graph Neural Network Based Adaptive Threat Detection for Cloud Identity and Access Management Logs

2025-12-11 · Venkata Tanuja Madireddy arxiv

The rapid expansion of cloud infrastructures and distributed identity systems has significantly increased the complexity and attack surface of modern enterprises. Traditional rule based or signature driven detection syst…

Graph Neural NetworkGraph Embedding