paper-with-me

Papers

TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment

2025-12-10 · Kanghyun Baek, Sangyub Lee, Jin Young Choi, Jaewoo Song, Daemin Park, Jooyoung Choi, Chaehun Shin, Bohyung Han, Sungroh Yoon arxiv

Despite recent advances, diffusion-based text-to-image models still struggle with accurate text rendering. Several studies have proposed fine-tuning or training-free refinement methods for accurate text rendering. However, the critical issue of text omission, where the desired text is partially or entirely missing, remains largely overlooked. In this work, we propose TextGuider, a novel training-free method that encourages accurate and complete text appearance by aligning textual content tokens and text regions in the image. Specifically, we analyze attention patterns in Multi-Modal Diffusion Transformer(MM-DiT) models, particularly for text-related tokens intended to be rendered in the image. Leveraging this observation, we apply latent guidance during the early stage of denoising steps based on two loss functions that we introduce. Our method achieves state-of-the-art performance in test-time text rendering, with significant gains in recall and strong results in OCR accuracy and CLIP score.

📄 PDF Abstract BibTeX arXiv:2512.09350

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models

2025-05-10 · Shuhan Zhuang, Mengqi Huang, Fengyi Fu, Nan Chen 외

Visual text rendering, which aims to accurately integrate specified textual content within generated images, is critical for various applications such as commercial design. Despite recent advances, current methods strugg…

Text Generation

Perceptual Similarity guidance and text guidance optimization for Editing Real Images using Guided Diffusion Models

2023-12-09 · Ruichen Zhang

When using a diffusion model for image editing, there are times when the modified image can differ greatly from the source. To address this, we apply a dual-guidance approach to maintain high fidelity to the original in …

Dynamic Classifier-Free Diffusion Guidance via Online Feedback

2025-09-19 · Pinelopi Papalampidi, Olivia Wiles, Ira Ktena, Aleksandar Shtedritski 외 arxiv

Classifier-free guidance (CFG) is a cornerstone of text-to-image diffusion models, yet its effectiveness is limited by the use of static guidance scales. This "one-size-fits-all" approach fails to adapt to the diverse re…

FlowMotion: Training-Free Flow Guidance for Video Motion Transfer

2026-03-06 · Zhen Wang, Youcan Xu, Jun Xiao, Long Chen arxiv

Video motion transfer aims to generate a target video that inherits motion patterns from a source video while rendering new scenes. Existing training-free approaches focus on constructing motion guidance based on the int…

Towards Training-Free Scene Text Editing

2026-03-25 · Yubo Li, Xugong Qin, Peng Zhang, Hailun Lin 외 arxiv

Scene text editing seeks to modify textual content in natural images while maintaining visual realism and semantic consistency. Existing methods often require task-specific training or paired data, limiting their scalabi…