paper-with-me

홈 › Papers

Evaluating Reasoning Fidelity in Visual Text Generation

2026-06-03 · Jiajun Hong, Jiawei Zhou arxiv

Recent text-to-image (T2I) models can render highly legible and well-structured text within images, enabling applications including document generation and slide generation. However, it remains unclear whether such systems faithfully preserve reasoning ability when complex solutions must be expressed directly through rendered text, or whether they merely imitate surface-level patterns. We investigate this question by evaluating reasoning fidelity in visual text generation, where models must express complete reasoning processes as images. Our evaluation includes long text rendering, factual knowledge probing, context understanding, and multi-step reasoning. Across these settings, we find that current T2I models frequently produce semantic errors, logical inconsistencies, and incorrect intermediate steps, even when the rendered text appears visually clear. These failures contrast with the strong reasoning performance of text-only models on the same tasks. Our findings reveal a substantial gap between visual text generation and procedural reasoning, motivating more reliable visual text reasoning.

📄 PDF Abstract BibTeX arXiv:2606.04479

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?

2026-02-27 · Hongbo Jiang, Jie Li, Yunhang Shen, Pingyang Dai 외 arxiv

Unified Multimodal Large Language Models (U-MLLMs) integrate understanding and generation within a single architecture. However, existing evaluations typically assess these capabilities separately, overlooking semantic e…

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

2026-06-26 · Anya Ji, Abhijith Varma Mudunuri, David M. Chan, Alane Suhr arxiv

While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains unclear whether they can recover temporal…

Visual ReasoningCode Generation

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

2026-06-11 · Tingyu Li, Le Zhou, Siyuan Li, Yujun Wu 외 arxiv

Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in complex long-chain scenarios, we identify a …

Reinforcement LearningImage Generation

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation

2026-05-07 · Yixin Zhu, Zixiong Wang, Jian Yang, Jin Xie 외 arxiv

Reliable simulation evaluation of robot manipulation policies serves as a high-fidelity proxy for real-world performance. Although existing benchmarks cover a wide range of task categories, they lack visual realism, crea…

Robot Manipulation

Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning

2025-12-10 · Xinyu Liu, Hangjie Yuan, Yujie Wei, Jiazheng Xing 외 arxiv

Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vision-language models (VLMs). We attribut…

Video Generation