paper-with-me

홈 › Papers

TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation

2024-04-18 · Tianyi Liang, Jiangqi Liu, Yifei HUANG, Shiqi Jiang, Jianshen Shi, Changbo Wang, Chenhui Li

Text-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate text placement without compromising image quality. This capability is non-trivial for real-world applications like graphic design, where clear visual hierarchy between content and text is essential. Prior work has primarily focused on arranging layouts within existing static images, leaving unexplored the potential of T2I models for generating text-friendly backgrounds. We present TextCenGen, a training-free dynamic background adaptation in the blank region for text-friendly image generation. Instead of directly reducing attention in text areas, which degrades image quality, we relocate conflicting objects before background optimization. Our method analyzes cross-attention maps to identify conflicting objects overlapping with text regions and uses a force-directed graph approach to guide their relocation, followed by attention excluding constraints to ensure smooth backgrounds. Our method is plug-and-play, requiring no additional training while well balancing both semantic fidelity and visual quality. Evaluated on our proposed text-friendly T2I benchmark of 27,000 images across four seed datasets, TextCenGen outperforms existing methods by achieving 23% lower saliency overlap in text regions while maintaining 98% of the semantic fidelity measured by CLIP score and our proposed Visual-Textual Concordance Metric (VTCM).

📄 PDF Abstract BibTeX arXiv:2404.11824

Code (1)

tianyilt/textcengen_background_adapt 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

2026-08-21 · Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat 외 arxiv

Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce sho…

Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning

2026-01-05 · Sungjune Park, Hongda Mao, Qingshuang Chen, Yong Man Ro 외 arxiv

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the i…

TEXTOC: Text-driven Object-Centric Style Transfer

2024-08-16 · Jihun Park, Jongmin Gim, Kyoungmin Lee, Seunghun Lee 외

We present Text-driven Object-Centric Style Transfer (TEXTOC), a novel method that guides style transfer at an object-centric level using textual inputs. The core of TEXTOC is our Patch-wise Co-Directional (PCD) loss, me…

ObjectStyle Transfer

Style-Editor: Text-driven Object-centric Style Editing

2025-01-01 · CVPR 2025 1 · Jihun Park, Jongmin Gim, Kyoungmin Lee, Seunghun Lee 외

We present Text-driven object-centric style editing model named Style-Editor, a novel method that guides style editing at an object-centric level using textual inputs.The core of Style-Editor is our Patch-wise Co-Dir…

Object

VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization

2026-03-07 · Seul-Ki Yeom, Marcel Simon, Eunbin Lee, Tae-Ho Kim arxiv

Self-supervised learning (SSL) has made rapid progress, yet learned features often over-rely on contextual shortcuts-background textures and co-occurrence statistics. While video provides rich temporal variation, dense i…

Self-Supervised Learning