paper-with-me

홈 › Papers

Taming Identity Consistency and Prompt Diversity in Diffusion Models via Latent Concatenation and Masked Conditional Flow Matching

2025-11-11 · Aditi Singhania, Arushi Jain, Krutik Malani, Riddhi Dhawan, Souymodip Chakraborty, Vineet Batra, Ankit Phogat arxiv

Subject-driven image generation aims to synthesize novel depictions of a specific subject across diverse contexts while preserving its core identity features. Achieving both strong identity consistency and high prompt diversity presents a fundamental trade-off. We propose a LoRA fine-tuned diffusion model employing a latent concatenation strategy, which jointly processes reference and target images, combined with a masked Conditional Flow Matching (CFM) objective. This approach enables robust identity preservation without architectural modifications. To facilitate large-scale training, we introduce a two-stage Distilled Data Curation Framework: the first stage leverages data restoration and VLM-based filtering to create a compact, high-quality seed dataset from diverse sources; the second stage utilizes these curated examples for parameter-efficient fine-tuning, thus scaling the generation capability across various subjects and contexts. Finally, for filtering and quality assessment, we present CHARIS, a fine-grained evaluation framework that performs attribute-level comparisons along five key axes: identity consistency, prompt adherence, region-wise color fidelity, visual quality, and transformation diversity.

📄 PDF Abstract BibTeX arXiv:2511.08061

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningImage Generation

Similar Papers 제목 키워드 기반

DeCorStory: Gram-Schmidt Prompt Embedding Decorrelation for Consistent Storytelling

2026-02-01 · Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris arxiv

Maintaining visual and semantic consistency across frames is a key challenge in text-to-image storytelling. Existing training-free methods, such as One-Prompt-One-Story, concatenate all prompts into a single sequence, wh…

Locate, Assign, Refine: Taming Customized Promptable Image Inpainting

2024-03-28 · Yulin Pan, Chaojie Mao, Zeyinzi Jiang, Zhen Han 외

Prior studies have made significant progress in image inpainting guided by either text description or subject image. However, the research on inpainting with flexible guidance or control, i.e., text-only, image-only, and…

Image Inpainting

ID-Booth: Identity-consistent Face Generation with Diffusion Models

2025-04-10 · Darian Tomašević, Fadi Boutros, Chenhao Lin, Naser Damer 외

Recent advances in generative modeling have enabled the generation of high-quality synthetic data that is applicable in a variety of domains, including face recognition. Here, state-of-the-art generative models typically…

DenoisingDiversityFace GenerationFace Recognition+3

CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation

2025-07-11 · Zhanxin Gao, Beier Zhu, Liang Yao, Jian Yang 외 arxiv

Subject-consistent generation (SCG)-aiming to maintain a consistent subject identity across diverse scenes-remains a challenge for text-to-image (T2I) models. Existing training-free SCG methods often achieve consistency …

Text-to-Image GenerationVisual Storytelling

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

2025-12-29 · Donghao Zhou, Jingyu Lin, Guibao Shen, Quande Liu 외 arxiv

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistenc…

Story Generation