paper-with-me

Papers

BLINK: Multimodal Large Language Models Can See but Not Perceive

2024-04-18 · Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning). However, we find these perception-demanding tasks cast significant challenges for current multimodal LLMs because they resist mediation through natural language. Blink reformats 14 classic computer vision tasks into 3,807 multiple-choice questions, paired with single or multiple images and visual prompting. While humans get 95.70% accuracy on average, Blink is surprisingly challenging for existing multimodal LLMs: even the best-performing GPT-4V and Gemini achieve accuracies of 51.26% and 45.72%, only 13.17% and 7.63% higher than random guessing, indicating that such perception abilities have not "emerged" yet in recent multimodal LLMs. Our analysis also highlights that specialist CV models could solve these problems much better, suggesting potential pathways for future improvements. We believe Blink will stimulate the community to help multimodal LLMs catch up with human-level visual perception.

📄 PDF Abstract BibTeX arXiv:2404.12390

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMultiple-choiceVisual Prompting

Similar Papers 제목 키워드 기반

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

2025-12-11 · Yuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, perceive complex scenes efficiently by dyna…

BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

2025-10-10 · Junyan Ye, Dongzhi Jiang, Jun He, Baichuan Zhou 외 arxiv

Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning,…

Visual Reasoning

mEBAL: A Multimodal Database for Eye Blink Detection and Attention Level Estimation

2020-06-09 · Roberto Daza, Aythami Morales, Julian Fierrez, Ruben Tolosana

This work presents mEBAL, a multimodal database for eye blink detection and attention level estimation. The eye blink frequency is related to the cognitive activity and automatic detectors of eye blinks have been propose…

EEGElectroencephalogram (EEG)Face Anti-Spoofing

MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

2025-08-04 · Mahtab Bigverdi, Wisdom Ikezogwo, Kevin Zhang, Hyewon Jeong 외 arxiv

Multimodal language models (MLMs) show promise for clinical decision support and diagnostic reasoning, raising the prospect of end-to-end automated medical image interpretation. However, clinicians are highly selective i…

Visual Grounding

mEBAL2 Database and Benchmark: Image-based Multispectral Eyeblink Detection

2023-09-14 · Roberto Daza, Aythami Morales, Julian Fierrez, Ruben Tolosana 외

This work introduces a new multispectral database and novel approaches for eyeblink detection in RGB and Near-Infrared (NIR) individual images. Our contributed dataset (mEBAL2, multimodal Eye Blink and Attention Level es…

EEGElectroencephalogram (EEG)Eyeblink detection