paper-with-me

Papers

Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

2024-03-14 · Zeyu Liu, Weicong Liang, Zhanhao Liang, Chong Luo, Ji Li, Gao Huang, Yuhui Yuan

Visual text rendering poses a fundamental challenge for contemporary text-to-image generation models, with the core problem lying in text encoder deficiencies. To achieve accurate text rendering, we identify two crucial requirements for text encoders: character awareness and alignment with glyphs. Our solution involves crafting a series of customized text encoder, Glyph-ByT5, by fine-tuning the character-aware ByT5 encoder using a meticulously curated paired glyph-text dataset. We present an effective method for integrating Glyph-ByT5 with SDXL, resulting in the creation of the Glyph-SDXL model for design image generation. This significantly enhances text rendering accuracy, improving it from less than $20\%$ to nearly $90\%$ on our design image benchmark. Noteworthy is Glyph-SDXL's newfound ability for text paragraph rendering, achieving high spelling accuracy for tens to hundreds of characters with automated multi-line layouts. Finally, through fine-tuning Glyph-SDXL with a small set of high-quality, photorealistic images featuring visual text, we showcase a substantial improvement in scene text rendering capabilities in open-domain real images. These compelling outcomes aim to encourage further exploration in designing customized text encoders for diverse and challenging tasks.

📄 PDF Abstract BibTeX arXiv:2403.09622

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering

2024-06-14 · Zeyu Liu, Weicong Liang, Yiming Zhao, Bohan Chen 외

Recently, Glyph-ByT5 has achieved highly accurate visual text rendering performance in graphic design images. However, it still focuses solely on English and performs relatively poorly in terms of visual appeal. In this …

GlyphControl: Glyph Conditional Control for Visual Text Generation

2023-05-29 · NeurIPS 2023 11 · Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang 외

Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient ap…

Optical Character Recognition (OCR)Text Generation

GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering

2026-03-16 · Xincheng Shuai, Ziye Li, Henghui Ding, Dacheng Tao arxiv

Generating accurate glyphs for visual text rendering is essential yet challenging. Existing methods typically enhance text rendering by training on a large amount of high-quality scene text images, but the limited covera…

Reinforcement Learning

HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models

2025-05-10 · Shuhan Zhuang, Mengqi Huang, Fengyi Fu, Nan Chen 외

Visual text rendering, which aims to accurately integrate specified textual content within generated images, is critical for various applications such as commercial design. Despite recent advances, current methods strugg…

Text Generation

GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing

2025-05-08 · CVPR 2025 1 · Tong Wang, Ting Liu, Xiaochao Qu, Chengjing Wu 외

Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promi…

Optical Character Recognition (OCR)Scene Text EditingText Generation