paper-with-me

Papers

Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models

2026-03-09 · Heng Zhou, Ao Yu, Li Kang, Yuchen Fan, Yutao Fan, Xiufeng Song, Hejia Geng, Yiran Qin arxiv

Vision-Language Models achieve near-perfect accuracy at reading text in images, yet prove largely typography-blind: capable of recognizing what text says, but not how it looks. We systematically investigate this gap by evaluating font family, size, style, and color recognition across 26 fonts, four scripts, and three difficulty levels. Our evaluation of 15 state-of-the-art VLMs reveals a striking perception hierarchy: color recognition is near-perfect, yet font style detection remains universally poor. We further find that model scale fails to predict performance and that accuracy is uniform across difficulty levels, together pointing to a training-data omission rather than a capacity ceiling. LoRA fine-tuning on a small set of synthetic samples substantially improves an open-source model, narrowing the gap to the best closed-source system and surpassing it on font size recognition. Font style alone remains resistant to fine-tuning, suggesting that relational visual reasoning may require architectural innovation beyond current patch-based encoders. We release our evaluation framework, data, and fine-tuning recipe to support progress in closing the typographic gap in vision-language understanding.

📄 PDF Abstract BibTeX arXiv:2603.08497

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

FonTS: Text Rendering with Typography and Style Controls

2024-11-28 · Wenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu 외

Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show prom…

parameter-efficient fine-tuning

The Cognitive Type Project -- Mapping Typography to Cognition

2024-03-06 · Nik Bear Brown

The Cognitive Type Project is focused on developing computational tools to enable the design of typefaces with varying cognitive properties. This initiative aims to empower typographers to craft fonts that enhance click-…

SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision

2026-01-24 · Jasmine Lesner, Michael Beyeler arxiv

Retinal prostheses restore limited visual perception, but low spatial resolution and temporal persistence make reading difficult. In sequential letter presentation, the afterimage of one symbol can interfere with percept…

Typography-Based Monocular Distance Estimation Framework for Vehicle Safety Systems

2026-03-24 · Manognya Lokesh Reddy, Zheng Liu arxiv

Accurate inter-vehicle distance estimation is a cornerstone of advanced driver assistance systems and autonomous driving. While LiDAR and radar provide high precision, their cost prohibits widespread adoption in mass-mar…

Autonomous Driving

Unsupervised Typography Transfer

2018-02-07 · Hanfei Sun, Yiming Luo, Ziang Lu

Traditional methods in Chinese typography synthesis view characters as an assembly of radicals and strokes, but they rely on manual definition of the key points, which is still time-costing. Some recent work on computer …