paper-with-me

Papers

InstanceControl: Controllable Complex Image Generation without Instance Labeling

2026-06-30 · Xiaoyu Liu, Huan Wang, Fan Li, Zhixin Wang, Jiaqi Xu, Ming Liu, Wangmeng Zuo hf

Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.

📄 PDF Abstract BibTeX arXiv:2606.31924

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

2026-06-29 · Qin Guo, Hao Luo, Dongxu Yue, Weixuan Jin 외 arxiv

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and …

Image Generation

IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts

2023-10-09 · Bohan Zeng, Shanglin Li, Yutang Feng, Ling Yang 외

Recent advances in 3D generation have been remarkable, with methods such as DreamFusion leveraging large-scale text-to-image diffusion-based models to guide 3D object generation. These methods enable the synthesis of det…

3D GenerationImage to 3DObjectText to 3D

Text2Street: Controllable Text-to-image Generation for Street Views

2024-02-07 · Jinming Su, Songen Gu, Yiting Duan, Xingyue Chen 외

Text-to-image generation has made remarkable progress with the emergence of diffusion models. However, it is still a difficult task to generate images for street views based on text, mainly because the road topology of s…

Image GenerationLayout GenerationObjectText to Image Generation+1

Layout Control and Semantic Guidance with Attention Loss Backward for T2I Diffusion Model

2024-11-11 · Guandong Li

Controllable image generation has always been one of the core demands in image generation, aiming to create images that are both creative and logical while satisfying additional specified conditions. In the post-AIGC era…

AttributeImage Generation

Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation

2025-01-01 · CVPR 2025 1 · Qitong Yang, Mingtao Feng, Zijie Wu, Weisheng Dong 외

3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex pre…

3D GenerationMamba