paper-with-me

Papers

MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

2024-10-17 · Donghao Zhou, Jiancheng Huang, Jinbin Bai, Jiaze Wang, Hao Chen, Guangyong Chen, Xiaowei Hu, Pheng-Ann Heng

Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. This task faces two challenges: semantic pollution, where undesired elements disrupt the target concept, and semantic imbalance, which causes disproportionate learning of the target concept and component. To address these, we design MagicTailor, a framework that uses Dynamic Masked Degradation to adaptively perturb unwanted visual semantics and Dual-Stream Balancing for more balanced learning of desired visual semantics. The experimental results show that MagicTailor achieves superior performance in this task and enables more personalized and creative image generation.

📄 PDF Abstract BibTeX arXiv:2410.13370

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

2025-10-21 · Chenghao Zhu, Meiling Tao, Tiannan Wang, Dongyi Ding 외 arxiv

Faithfully personalizing large language models (LLMs) to align with individual user preferences is a critical but challenging task. While supervised fine-tuning (SFT) quickly reaches a performance plateau, standard reinf…

Reinforcement Learning

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models

2026-05-19 · Junhao Li, Xinhao Zhong, Yi sun, Yuxia Qiao 외 arxiv

Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation. Despite their strong generative capability, existing VAR-based personalization methods remain limited to sta…

Text-to-Image Generation

TextBoost: Towards One-Shot Personalization of Text-to-Image Models via Fine-tuning Text Encoder

2024-09-12 · Nahyeon Park, Kunhee Kim, Hyunjung Shim

Recent breakthroughs in text-to-image models have opened up promising research avenues in personalized image generation, enabling users to create diverse images of a specific subject using natural language prompts. Howev…

Diffusion PersonalizationDisentanglementImage GenerationPersonalized Image Generation+1

Unified Personalized Understanding, Generating and Editing

2026-01-11 · Yu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan 외 arxiv

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to mode…

Image Editing

Reverse Personalization

2025-12-28 · Han-Wei Kung, Tuomas Varanka, Nicu Sebe arxiv

Recent text-to-image diffusion models have demonstrated remarkable generation of realistic facial images conditioned on textual prompts and human identities, enabling creating personalized facial imagery. However, existi…

Face Anonymization