paper-with-me

홈 › Papers

ReViCo: Unveiling the Limitations of VLMs in Visual Text Understanding via Error Correction

2026-08-27 · Bojun Zhang, Junhong Liang, Feifei Zhai, Fengxian Ji, Yu Zhou arxiv

Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.

📄 PDF Abstract BibTeX arXiv:2608.27154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision Language Models Are Not (Yet) Spelling Correctors

2025-09-22 · Junhong Liang, Bojun Zhang arxiv

Spelling correction from visual input poses unique challenges for vision language models (VLMs), as it requires not only detecting but also correcting textual errors directly within images. We present ReViCo (Real Visual…

When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI

2025-11-29 · Yanhui Li, Qi Zhou, Zhihong Xu, Huizhong Guo 외 arxiv

Large vision-language models (LVLMs) are increasingly used for tasks where detecting multimodal harmful content is crucial, such as online content moderation. However, real-world harmful content is often camouflaged, rel…

Scene UnderstandingVisual Reasoning

Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens

2025-09-03 · Sohee Kim, Soohyun Ryu, Joonhyung Park, Eunho Yang arxiv

Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. However, our finding reveals they often mistakenly perceive text inputs lacking visual evid…

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

2025-05-26 · Tianle Li, Jihai Zhang, Yongming Rao, Yu Cheng

While large language models (LLMs) demonstrate strong reasoning capabilities utilizing reinforcement learning (RL) with verifiable reward, whether large vision-language models (VLMs) can directly inherit such capabilitie…

DiagnosticReinforcement Learning (RL)Visual Grounding

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

2026-05-26 · Xiang Fang, Wanlong Fang, Changshuo Wang arxiv

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness …

Visual Question AnsweringAutonomous DrivingImage Captioning