paper-with-me

Papers

FontFusion: Enhancing Generative Text in Diffusion Models with Typographic Conditioning

2026-06-04 · Marian Lupascu, Nipun Jindal, Ionut Mironica, Zhaowen Wang arxiv

Typography generation in diffusion models faces a persistent trade-off: enabling precise font control typically degrades text legibility, while maintaining readability often sacrifices typographic fidelity. We present FontFusion, a plug-and-play conditioning framework for Diffusion Transformer (DiT) architectures that resolves this dilemma through three core innovations: (1) a hierarchical token representation establishing explicit text-font relationships at multiple granularities, (2) position-aware embeddings creating spatial bindings between typography and image content, and (3) a multi-level token dropping strategy improving both computational efficiency and generalization to unseen fonts. Our systematic evaluation of font embedding spaces reveals that a dual encoder combining DeepFont and DINOv2 outperforms any single encoder for typography tasks. FontFusion demonstrates 76% relative improvement on challenging decorative fonts over single-encoder baselines and font consistency gains exceeding approximately 68-76% over unconditioned models, while integrating into existing DiT architectures without retraining.

📄 PDF Abstract BibTeX arXiv:2606.06066

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Typographic Text Generation with Off-the-Shelf Diffusion Model

2024-02-22 · KhayTze Peong, Seiichi Uchida, Daichi Haraguchi

Recent diffusion-based generative models show promise in their ability to generate text images, but limitations in specifying the styles of the generated texts render them insufficient in the realm of typographic design.…

Text Generation

FonTS: Text Rendering with Typography and Style Controls

2024-11-28 · Wenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu 외

Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show prom…

parameter-efficient fine-tuning

PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation

2026-03-25 · Yuheng Feng, Wen Zhang, Haodong Duan, Xingxing Zou arxiv

We present PosterIQ, a design-driven benchmark for poster understanding and generation, annotated across composition structure, typographic hierarchy, and semantic intent. It includes 7,765 image-annotation instances and…

Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

2025-08-28 · Lorenz Hufe, Constantin Venhoff, Erblina Purelku, Maximilian Dreyer 외 arxiv

Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how …

Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model

2024-02-29 · Hao Cheng, Erjia Xiao, Jindong Gu, Le Yang 외

Large Vision-Language Models (LVLMs) rely on vision encoders and Large Language Models (LLMs) to exhibit remarkable capabilities on various multi-modal tasks in the joint space of vision and language. However, the Typogr…

Language ModelingLanguage ModellingObject RecognitionZero-Shot Learning