paper-with-me

Papers

Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation

2026-04-28 · Hao Li arxiv

Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows that the \emph{starting noise} of a diffusion model carries significant semantic information: ``golden'' noise predicted from text can substantially raise prompt fidelity. We observe that this noise prediction is, however, fundamentally global: the same network is asked to summarise a long, multi-region prompt with a single text embedding, which becomes the bottleneck whenever the prompt describes scenes with spatially-separated entities. We introduce \textbf{Golden RPG}, a region-aware noise predictor that extends a frozen NPNet with two trainable additions: (i) a per-region \textbf{FiLM adapter} that reshapes the predicted noise according to each sub-prompt; and (ii) a \textbf{Region Cross-Attention} layer injected between two stages of the Swin backbone, allowing different spatial locations to attend to different sub-prompt tokens. To prevent the regional conditioning from degrading samples whose prompts are already easy, we further propose a \textbf{Confidence-Adaptive Blending} head that dynamically predicts, per sample, how strongly the regional signal should override the global signal. We evaluate on the original RPG benchmark (20 prompts, 100 samples) and on four multi-region categories of T2I-CompBench (1{,}200 images, six competing methods). Golden RPG achieves the highest Cross-Region-Coherence score on every category, while matching the strongest baselines on absolute CLIP-Score and CLIP-IQA. A paired user study further shows a $\boldsymbol{\sim}$67\% preference over the strongest baseline. The adapter contains $\sim$2M trainable parameters and adds only $0.6$\,s of inference overhead on top of SDXL.

📄 PDF Abstract BibTeX arXiv:2604.25314

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Golden Noise for Diffusion Models: A Learning Framework

2024-11-14 · Zikai Zhou, Shitong Shao, Lichen Bai, Zhiqiang Xu 외

Text-to-image diffusion model is a popular paradigm that synthesizes personalized images by providing a text prompt and a random Gaussian noise. While people observe that some noises are ``golden noises'' that can achiev…

Prompt Learning

CAD: Confidence-Aware Adaptive Displacement for Semi-Supervised Medical Image Segmentation

2025-02-01 · Wenbo Xiao, Zhihao Xu, Guiping Liang, Yangjun Deng 외

Semi-supervised medical image segmentation aims to leverage minimal expert annotations, yet remains confronted by challenges in maintaining high-quality consistency learning. Excessive perturbations can degrade alignment…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1

Robust Partial 3D Point Cloud Registration via Confidence Estimation under Global Context

2025-09-29 · Yongqiang Wang, Weigang Li, Wenping Liu, Zhe Xu 외 arxiv

Partial point cloud registration is essential for autonomous perception and 3D scene understanding, yet it remains challenging owing to structural ambiguity, partial visibility, and noise. We address these issues by prop…

Point Cloud RegistrationScene Understanding

CASC-AI: Consensus-aware Self-corrective AI Agents for Noise Cell Segmentation

2025-02-11 · Ruining Deng, Yihe Yang, David J. Pisapia, Benjamin Liechty 외

Multi-class cell segmentation in high-resolution gigapixel whole slide images (WSI) is crucial for various clinical applications. However, training such models typically requires labor-intensive, pixel-wise annotations b…

AI AgentCell SegmentationContrastive Learningwhole slide images

Inference for Heteroskedastic PCA with Missing Data

2021-07-26 · Yuling Yan, Yuxin Chen, Jianqing Fan

This paper studies how to construct confidence regions for principal component analysis (PCA) in high dimension, a problem that has been vastly under-explored. While computing measures of uncertainty for nonlinear/noncon…

valid