paper-with-me

홈 › Papers

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

2025-12-11 · Yuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen, Peng Fu, Zheng Lin, Shuohuan Wang, Yu Sun, Hua Wu, Weiping Wang, Haifeng Wang arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, perceive complex scenes efficiently by dynamically scanning and focusing on salient regions in a sequential "blink-like" process. Motivated by this strategy, we first investigate whether MLLMs exhibit similar behavior. Our pilot analysis reveals that MLLMs naturally attend to different visual regions across layers and that selectively allocating more computation to salient tokens can enhance visual perception. Building on this insight, we propose Blink, a dynamic visual token resolution framework that emulates the human-inspired process within a single forward pass. Specifically, Blink includes two modules: saliency-guided scanning and dynamic token resolution. It first estimates the saliency of visual tokens in each layer based on the attention map, and extends important tokens through a plug-and-play token super-resolution (TokenSR) module. In the next layer, it drops the extended tokens when they lose focus. This dynamic mechanism balances broad exploration and fine-grained focus, thereby enhancing visual perception adaptively and efficiently. Extensive experiments validate Blink, demonstrating its effectiveness in enhancing visual perception and multimodal understanding.

📄 PDF Abstract BibTeX arXiv:2512.10548

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Real-Time Face & Eye Tracking and Blink Detection using Event Cameras

2020-10-16 · Cian Ryan, Brian O Sullivan, Amr Elrasad, Joe Lemley 외

Event cameras contain emerging, neuromorphic vision sensors that capture local light intensity changes at each pixel, generating a stream of asynchronous events. This way of acquiring visual information constitutes a dep…

Event-based Face Detection and Tracking in the Blink of an Eye

2018-03-27 · Gregor Lenz, Sio-Hoi Ieng, Ryad Benosman

We present the first purely event-based method for face detection using the high temporal resolution of an event-based camera. We will rely on a new feature that has never been used for such a task that relies on detecti…

Face DetectionPosition

Superresolution imaging of single DNA molecules using stochastic photoblinking of minor groove and intercalating dyes

2015-04-14

As proof-of-principle for generating superresolution structural information from DNA we applied a method of localization microscopy utilizing photoblinking comparing intercalating dye YOYO-1 against minor groove binding …

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

2026-06-29 · Xiaoqian Shen, Mohamed Elhoseiny arxiv

Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images and long-context videos remains challenging. As the number of visual…

BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception

2025-10-10 · Junyan Ye, Dongzhi Jiang, Jun He, Baichuan Zhou 외 arxiv

Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning,…

Visual Reasoning