paper-with-me

Papers

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

2026-06-05 · Lujun Li, Lama Sleem, Niccolo Gentile, Yangjie Xu, Yewei Song, Wenbo Wu, Radu State arxiv

Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored. A natural extension of ``How many r are there in Strawberry?'' asks: how small a visual pattern can a VLM reliably perceive? As such, we introduce FineSightBench, a new benchmark that systematically probes this limit by separating perception tasks (pixel-level recognition of letters, shapes, objects) from reasoning tasks (spatial reasoning, counting, ordering over small targets) across controlled scales of 4--48px. Through comprehensive experiments and detailed failure mode analysis on state-of-the-art models, we reveal a sharp dissociation: perception saturates around 12px, while reasoning remains limited even at larger scales, with persistent numeracy and sequence errors. These findings expose fundamental deficiencies in VLMs' fine-scale visual reasoning that demand more rigorous evaluation.

📄 PDF Abstract BibTeX arXiv:2606.07861

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

Everything that can be learned about a causal structure with latent variables by observational and interventional probing schemes

2024-07-01 · Marina Maciel Ansanelli, Elie Wolfe, Robert W. Spekkens

What types of differences among causal structures with latent variables are impossible to distinguish by statistical data obtained by probing each visible variable? If the probing scheme is simply passive observation, th…

XoFTR: Cross-modal Feature Matching Transformer

2024-04-15 · Önder Tuzcuoğlu, Aybora Köksal, Buğra Sofu, Sinan Kalkan 외

We introduce, XoFTR, a cross-modal cross-view method for local feature matching between thermal infrared (TIR) and visible images. Unlike visible images, TIR images are less susceptible to adverse lighting and weather co…

Image Augmentation

Adaptive Assignment for Geometry Aware Local Feature Matching

2022-07-18 · CVPR 2023 1 · Dihe Huang, Ying Chen, Shang Xu, Yong liu 외

The detector-free feature matching approaches are currently attracting great attention thanks to their excellent performance. However, these methods still struggle at large-scale and viewpoint variations, due to the geom…

Feature Correlation

Generative adversarial network based single pixel imaging

2021-07-11 · Ming Zhao, Fengqiang Li, Fengyue Huo, Zhiming Tian

Single pixel imaging can reconstruct two-dimensional images of a scene with only a single-pixel detector. It has been widely used for imaging in non-visible bandwidth (e.g., near-infrared and X-ray) where focal-plane arr…

Generative Adversarial Network

A Synthesis-Based Approach for Thermal-to-Visible Face Verification

2021-08-21 · Neehar Peri, Joshua Gleason, Carlos D. Castillo, Thirimachos Bourlai 외

In recent years, visible-spectrum face verification systems have been shown to match the performance of experienced forensic examiners. However, such systems are ineffective in low-light and nighttime conditions. Thermal…

Face AlignmentFace GenerationFace Verification