paper-with-me

홈 › Papers

Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers

2026-03-06 · Ruidong Chen, Yancheng Bai, Xuanpu Zhang, Jianhao Zeng, Lanjun Wang, Dan Song, Lei Sun, Xiangxiang Chu, Anan Liu arxiv

Region-instructed layout control in text-to-image generation is highly practical, yet existing methods suffer from limitations: (i) training-based approaches inherit data bias and often degrade image quality, and (ii) current techniques struggle with occlusion order, limiting real-world usability. To address these issues, we propose LayerBind. By modeling regional generation as distinct layers and binding them during the generation, our method enables precise regional and occlusion controllability. Our motivation stems from the observation that spatial layout and occlusion are established at a very early denoising stage, suggesting that rearranging the early latent structure is sufficient to modify the final output. Building on this, we structure the scheme into two phases: instance initialization and subsequent semantic nursing. (1) First, leveraging the contextual sharing mechanism in multimodal joint attention, Layer-wise Instance Initialization creates per-instance branches that attend to their own regions while anchoring to the shared background. At a designated early step, these branches are fused according to the layer order to form a unified latent with a pre-established layout. (2) Then, Layer-wise Semantic Nursing reinforces regional details and maintains the occlusion order via a layer-wise attention enhancement. Specifically, a sequential layered attention path operates alongside the standard global path, with updates composited under a layer-transparency scheduler. LayerBind is training-free and plug-and-play, serving as a regional and occlusion controller across Diffusion Transformers. Beyond generation, it natively supports editable workflows, allowing for flexible modifications like changing instances or rearranging visible orders. Both qualitative and quantitative results demonstrate LayerBind's effectiveness, highlighting its strong potential for creative applications.

📄 PDF Abstract BibTeX arXiv:2603.05769

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Coupled Network for Robust Pedestrian Detection with Gated Multi-Layer Feature Extraction and Deformable Occlusion Handling

2019-12-18 · Tianrui Liu, Wenhan Luo, Lin Ma, Jun-Jie Huang 외

Pedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, detecting small-scaled pedestrians and occluded pedestrians remains a challenging pr…

Occlusion HandlingPedestrian Detection

Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

2024-11-10 · Zhennan Chen, Yajie Li, Haofan Wang, Zhibo Chen 외

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce add…

AttributeImage GenerationRAGText to Image Generation+1

Optimizing Vision-Language Consistency via Cross-Layer Regional Attention Alignment

2025-07-31 · Yifan Wang, Hongfeng Ai, Quangao Liu, Maowei Jiang 외 arxiv

Vision Language Models (VLMs) face challenges in effectively coordinating diverse attention mechanisms for cross-modal embedding learning, leading to mismatched attention and suboptimal performance. We propose Consistent…

Gated Multi-layer Convolutional Feature Extraction Network for Robust Pedestrian Detection

2019-10-25 · Tianrui Liu, Jun-Jie Huang, Tianhong Dai, Guangyu Ren 외

Pedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, robustly detecting pedestrians with a large variant on sizes and with occlusions rem…

Pedestrian Detection

Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers

2021-03-23 · CVPR 2021 1 · Lei Ke, Yu-Wing Tai, Chi-Keung Tang

Segmenting highly-overlapping objects is challenging, because typically no distinction is made between real object contours and occlusion boundaries. Unlike previous two-stage instance segmentation methods, we model imag…

Amodal Instance SegmentationBoundary DetectionInstance SegmentationObject Detection+3