paper-with-me

홈 › Papers

MS-CustomNet: Controllable Multi-Subject Customization with Hierarchical Relational Semantics

2026-03-22 · Pengxiang Cai, Mengyang Li arxiv

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle to provide explicit user-defined control over the compositional structure and precise spatial relationships between subjects. To address this, we introduce MS-CustomNet, a novel framework for multi-subject customization. MS-CustomNet allows zero-shot integration of multiple user-provided objects and, crucially, empowers users to explicitly define these hierarchical arrangements and spatial placements within the generated image. Our approach ensures individual subject identity preservation while learning and enacting these user-specified inter-subject compositions. We also present the MSI dataset, derived from COCO, to facilitate training on such complex multi-subject compositions. MS-CustomNet offers enhanced, fine-grained control over multi-subject image generation. Our method achieves a DINO-I score of 0.61 for identity preservation and a YOLO-L score of 0.94 for positional control in multi-subject customization tasks, demonstrating its superior capability in generating high-fidelity images with precise, user-directed multi-subject compositions and spatial control.

📄 PDF Abstract BibTeX arXiv:2603.21136

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models

2023-10-30 · Ziyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi 외

Incorporating a customized object into image generation presents an attractive feature in text-to-image generation. However, existing optimization-based and encoder-based methods are hindered by drawbacks such as time-co…

Image GenerationNovel View SynthesisObjectText to Image Generation+1

PositionIC: Unified Position and Identity Consistency for Image Customization

2025-07-18 · Junjie Hu, Tianyang Han, Kai Ma, Jialin Gao 외 arxiv

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a sca…

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

2026-03-12 · Yujie Wei, Xinyu Liu, Shiwei Zhang, Hangjie Yuan 외 arxiv

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge t…

Reinforcement Learning

MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion

2025-10-15 · Minjung Shin, Hyunin Cho, Sooyeon Go, Jin-Hwa Kim 외 arxiv

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support custom…

Cones 2: Customizable Image Synthesis with Multiple Subjects

2023-05-30 · Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng 외

Synthesizing images with user-specified subjects has received growing attention due to its practical applications. Despite the recent success in single subject customization, existing algorithms suffer from high training…

Image Generation