paper-with-me

홈 › Papers

Training-Free Occluded Text Rendering via Glyph Priors and Attention-Guided Semantic Blending

2026-05-16 · Jingqi Hou, Hongtian Wang arxiv

We present a training-free framework for occluded text rendering with a pretrained FLUX.1-dev backbone. The task requires a model to render recognizable typography and place an occluding object over the intended text region. This setting remains difficult for existing text-to-image generators: the occluder often drifts away from the text, while the text may be distorted or appear to float on top of the occluding object. To address this problem, we propose a restarted dual-stream inference framework that decouples text-layout preservation from occluder insertion. A Base Stream provides a clean typographic reference and same-step key/value (K/V) features, while the Edit Stream is conditioned on the occlusion prompt. We further adopt the spectral glyph-prior idea from FreeText and adapt it to stabilize the target text structure during early-to-mid denoising. In the reasoning pass, our method localizes the target text, estimates a text-band region from token-conditioned attention and glyph support, and derives an anchor-aware hard fusion mask for the occluder. In the final edit pass, generation restarts from the same initial noise and applies hard mask-guided image-token K/V replacement at selected attention sites, preserving the Base layout outside the mask while injecting the occluder appearance from the Edit Stream inside the mask. Experiments on representative occluded text scenarios demonstrate substantially improved text readability and competitive occlusion alignment, yielding more stable object-on-text compositions without any model fine-tuning.

📄 PDF Abstract BibTeX arXiv:2605.16810

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models

2025-05-10 · Shuhan Zhuang, Mengqi Huang, Fengyi Fu, Nan Chen 외

Visual text rendering, which aims to accurately integrate specified textual content within generated images, is critical for various applications such as commercial design. Despite recent advances, current methods strugg…

Text Generation

FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection

2026-01-02 · Ruiqiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang 외 arxiv

Large-scale text-to-image (T2I) diffusion models excel at open-domain synthesis but still struggle with precise text rendering, especially for multi-line layouts, dense typography, and long-tailed scripts such as Chinese…

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

2026-03-12 · Zexuan Yan, Jiarui Jin, Yue Ma, Shijian Wang 외 arxiv

Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems fr…

GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

2026-03-16 · Xincheng Shuai, Ziye Li, Henghui Ding, Dacheng Tao arxiv

Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality scene text images, but the limited covera…

Reinforcement Learning

First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending

2024-10-14 · Zhenhang Li, Yan Shu, Weichao Zeng, Dongbao Yang 외

Diffusion models, known for their impressive image generation abilities, have played a pivotal role in the rise of visual text generation. Nevertheless, existing visual text generation methods often focus on generating e…

Image GenerationText Generation