paper-with-me

홈 › Papers

Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

2024-11-10 · Zhennan Chen, Yajie Li, Haofan Wang, Zhibo Chen, Zhengkai Jiang, Jun Li, Qian Wang, Jian Yang, Ying Tai

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce additional trainable modules, thus only applicable to specific models, or manipulate on score maps within cross-attention layers using attention masks, resulting in limited control strength when the number of regions increases. To handle these limitations, we present RAG, a Regional-Aware text-to-image Generation method conditioned on regional descriptions for precise layout composition. RAG decouple the multi-region generation into two sub-tasks, the construction of individual region (Regional Hard Binding) that ensures the regional prompt is properly executed, and the overall detail refinement (Regional Soft Refinement) over regions that dismiss the visual boundaries and enhance adjacent interactions. Furthermore, RAG novelly makes repainting feasible, where users can modify specific unsatisfied regions in the last generation while keeping all other regions unchanged, without relying on additional inpainting models. Our approach is tuning-free and applicable to other frameworks as an enhancement to the prompt following property. Quantitative and qualitative experiments demonstrate that RAG achieves superior performance over attribute binding and object relationship than previous tuning-free methods.

📄 PDF Abstract BibTeX arXiv:2411.06558

Code (1)

nju-pcalab/rag-diffusion 공식 구현 pytorch

Tasks

AttributeImage GenerationRAGText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
WordPiece 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Hard-aware Instance Adaptive Self-training for Unsupervised Cross-domain Semantic Segmentation

2023-02-14 · Chuang Zhu, Kebin Liu, Wenqi Tang, Ke Mei 외

The divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to solve such problem. Recent works show t…

Domain AdaptationPseudo LabelSemantic SegmentationSynthetic-to-Real Translation+2

RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection

2024-05-30 · Fangyi Chen, Han Zhang, Zhantao Yang, Hao Chen 외

Open-vocabulary object detection (OVD) requires solid modeling of the region-semantic relationship, which could be learned from massive region-text pairs. However, such data is limited in practice due to significant anno…

Image CaptioningImage InpaintingObjectobject-detection+4

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

2026-05-13 · Amirreza Mohseni, Mona Mohammadi, Morteza Saghafian, Naser Talebizadeh Sardari arxiv

We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of \textbf{756 images} of pairwise non-intersecting Jordan curves across easy, polygonal, topographic-ins…

Structured PredictionVisual Reasoning

R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

2023-10-13 · Jiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang 외

Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified …

Image GenerationText to Image GenerationText-to-Image Generation

RegionE: Adaptive Region-Aware Generation for Efficient Image Editing

2025-10-29 · Pengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen 외 arxiv

Recently, instruction-based image editing (IIE) has received widespread attention. In practice, IIE often modifies only specific regions of an image, while the remaining areas largely remain unchanged. Although these two…

Image Editing