paper-with-me

홈 › Papers

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

2025-09-08 · Yufeng Cheng, Wenxu Wu, Shaojin Wu, Mengqi Huang, Fei Ding, Qian He arxiv

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-reference images, limiting the identity scalability of customization models. To address this, we present UMO, a Unified Multi-identity Optimization framework, designed to maintain high-fidelity identity preservation and alleviate identity confusion with scalability. With "multi-to-multi matching" paradigm, UMO reformulates multi-identity generation as a global assignment optimization problem and unleashes multi-identity consistency for existing image customization methods generally through reinforcement learning on diffusion models. To facilitate the training of UMO, we develop a scalable customization dataset with multi-reference images, consisting of both synthesised and real parts. Additionally, we propose a new metric to measure identity confusion. Extensive experiments demonstrate that UMO not only improves identity consistency significantly, but also reduces identity confusion on several image customization methods, setting a new state-of-the-art among open-source methods along the dimension of identity preserving. Code and model: https://github.com/bytedance/UMO

📄 PDF Abstract BibTeX arXiv:2509.06818

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

PositionIC: Unified Position and Identity Consistency for Image Customization

2025-07-18 · Junjie Hu, Tianyang Han, Kai Ma, Jialin Gao 외 arxiv

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a sca…

MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning

2025-10-31 · Hengjia Li, Jianjin Xu, Keli Cheng, Lei Wang 외 arxiv

Recent advances in personalized generative models have demonstrated impressive capabilities in producing identity-consistent images of the same individual across diverse scenes. However, most existing methods lack explic…

Semantic correspondence

Subject-driven Video Generation via Disentangled Identity and Motion

2025-04-23 · Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho 외

We propose to train a subject-driven customized video generation model through decoupling the subject-specific learning from temporal dynamics in zero-shot without additional tuning. A traditional method for video custom…

Subject-driven Video GenerationVideo Generation

IC-Custom: Diverse Image Customization via In-Context Learning

2025-07-02 · Yaowei Li, Xiaoyu Li, Zhaoyang Zhang, Yuxuan Bian 외 arxiv

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into po…

Customized Visual Storytelling with Unified Multimodal LLMs

2026-03-29 · Wei-Hua Li, Cheng Sun, Chu-Song Chen arxiv

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …

Visual StorytellingStory Generation