paper-with-me

Papers

Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models

2025-04-02 · Zhaochen Wang, Yujun Cai, Zi Huang, Bryan Hooi, Yiwei Wang, Ming-Hsuan Yang

Vision-language models (VLMs) have advanced rapidly in processing multimodal information, but their ability to reconcile conflicting signals across modalities remains underexplored. This work investigates how VLMs process ASCII art, a unique medium where textual elements collectively form visual patterns, potentially creating semantic-visual conflicts. We introduce a novel evaluation framework that systematically challenges five state-of-the-art models (including GPT-4o, Claude, and Gemini) using adversarial ASCII art, where character-level semantics deliberately contradict global visual patterns. Our experiments reveal a strong text-priority bias: VLMs consistently prioritize textual information over visual patterns, with visual recognition ability declining dramatically as semantic complexity increases. Various mitigation attempts through visual parameter tuning and prompt engineering yielded only modest improvements, suggesting that this limitation requires architectural-level solutions. These findings uncover fundamental flaws in how current VLMs integrate multimodal information, providing important guidance for future model development while highlighting significant implications for content moderation systems vulnerable to adversarial examples.

📄 PDF Abstract BibTeX arXiv:2504.01589

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Riemannian Geometry Speaks Louder Than Words: From Graph Foundation Model to Next-Generation Graph Intelligence

2026-03-23 · Philip S. Yu, Li Sun arxiv

Graphs provide a natural description of the complex relationships among objects, and play a pivotal role in communications, transportation, social computing, the life sciences, etc. Currently, there is strong agreement t…

Graph Learning

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

2025-12-02 · Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik 외 arxiv

Large language models (LLMs) have demonstrated several emergent behaviors with scale, including reasoning and fluency in long-form text generation. However, they continue to struggle with tasks requiring precise spatial …

Text Generation

ASCII Art Turns LLMs into VLA Controllers

2026-06-19 · Yitao Jiang, Roy Xing, Luyang Zhao, Brian Plancher 외 arxiv

Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal backbones with large data and compute requirements. We demonstrate tha…

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

2026-08-10 · Bo Wang, Ruixing Zhang, Yunqi Liu, Yang Zhang 외 hf

User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support mult…

Reinforcement Learning

Language Richness of the Web

2012-05-01 · LREC 2012 5 · Martin Majli{\v{s}}, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}

We have built a corpus containing texts in 106 languages from texts available on the Internet and on Wikipedia. The W2C Web Corpus contains 54.7{\textasciitilde}GB of text and the W2C Wiki Corpus contains 8.5{\textasciit…