paper-with-me

Papers

Vision language models are blind: Failing to translate detailed visual features into words

2024-07-09 · Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri, Anh Totti Nguyen

While large language models with vision capabilities (VLMs), e.g., GPT-4o and Gemini 1.5 Pro, score high on many vision-understanding benchmarks, they are still struggling with low-level vision tasks that are easy to humans. Specifically, on BlindTest, our suite of 7 very simple tasks, including identifying (a) whether two circles overlap; (b) how many times two lines intersect; (c) which letter is being circled in a word; and (d) the number of circles in an Olympic-like logo, four state-of-the-art VLMs are only 58.07% accurate on average. Claude 3.5 Sonnet performs the best at 77.84% accuracy, far from the human expected accuracy of 100%. Across different image resolutions and line widths, VLMs including slow-thinking models consistently struggle with those tasks that require precise spatial information when geometric primitives overlap or are close. Yet, VLMs perform at near-100% accuracy when much more space is added to separate shapes and letters. Linear probing experiments show that vision encoders contain sufficient visual information to solve BlindTest and that language models fail to decode this information into correct answers. Code and data are at: https://vlmsareblind.github.io

📄 PDF Abstract BibTeX arXiv:2407.06581

Code (1)

anguyen8/vision-llms-are-blind 공식 구현

Similar Papers 제목 키워드 기반

How Far Are We from Intelligent Visual Deductive Reasoning?

2024-03-07 · Yizhe Zhang, He Bai, Ruixiang Zhang, Jiatao Gu 외

Vision-Language Models (VLMs) have recently demonstrated incredible strides on diverse vision language tasks. We dig into vision-based deductive reasoning, a more sophisticated but less explored realm, and find previousl…

In-Context LearningVisual Reasoning

A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction

2023-10-31 · Yu Hao, Fan Yang, Hao Huang, Shuaihang Yuan 외

People with blindness and low vision (pBLV) encounter substantial challenges when it comes to comprehensive scene recognition and precise object identification in unfamiliar environments. Additionally, due to the vision …

Language ModelingLanguage ModellingPrompt EngineeringScene Recognition

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

2026-04-11 · Jingru Li, Wei Ren, Tianqing Zhu arxiv

Large Vision-Language Models (LVLMs) rely on attention-based retrieval of safety instructions to maintain alignment during generation. Existing attacks typically optimize image perturbations to maximize harmful output li…

Linguistic Blind Spots of Large Language Models

2025-03-25 · Jiali Cheng, Hadi Amiri

Large language models (LLMs) are the foundation of many AI applications today. However, despite their remarkable proficiency in generating coherent text, questions linger regarding their ability to perform fine-grained l…

Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation

2024-08-23 · Md Touhidul Islam, Imran Kabir, Elena Ariel Pearce, Md Alimoor Reza 외

This paper presents a curated list of 90 objects essential for the navigation of blind and low-vision (BLV) individuals, encompassing road, sidewalk, and indoor environments. We develop the initial list by analyzing 21 p…

Object