paper-with-me

홈 › Papers

VB: Visibility Benchmark for Visibility and Perspective Reasoning in Images

2026-03-03 · Neil Tripathi arxiv

We present VB, a benchmark that tests whether vision-language models can determine what is and is not visible in a photograph, and abstain when a human viewer cannot reliably answer. Each item pairs a single photo with a short yes/no visibility claim; the model must output VISIBLY_TRUE, VISIBLY_FALSE, or ABSTAIN, together with a confidence score. Items are organized into 100 families using a 2x2 design that crosses a minimal image edit with a minimal text edit, yielding 300 headline evaluation cells. Unlike prior unanswerable-VQA benchmarks, VB tests not only whether a question is unanswerable but why (via reason codes tied to specific visibility factors), and uses controlled minimal edits to verify that model judgments change when and only when the underlying evidence changes. We score models on confidence-aware accuracy with abstention (CAA), minimal-edit flip rate (MEFR), confidence-ranked selective prediction (SelRank), and second-order perspective reasoning (ToMAcc); all headline numbers are computed on the strict XOR subset (three cells per family, 300 scored items per model). We evaluate nine models spanning flagship and prior-generation closed-source systems, and open-source models from 8B to 12B parameters. GPT-4o and Gemini 3.1 Pro effectively tie for the best composite score (0.728 and 0.727), followed by Gemini 2.5 Pro (0.678). The best open-source model, Gemma 3 12B (0.505), surpasses one prior-generation closed-source system. Text-flip robustness exceeds image-flip robustness for six of nine models, and confidence calibration varies substantially: GPT-4o and Gemini 2.5 Pro achieve similar accuracy yet differ sharply in selective prediction quality.

📄 PDF Abstract BibTeX arXiv:2603.06680

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

2026-07-01 · Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li 외 arxiv

Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene…

Spatial Reasoning

Modeling Mutual Visibility Relationship in Pedestrian Detection

2013-06-01 · CVPR 2013 6 · Wanli Ouyang, Xingyu Zeng, Xiaogang Wang

Detecting pedestrians in cluttered scenes is a challenging problem in computer vision. The difficulty is added when several pedestrians overlap in images and occlude each other. We observe, however, that the occlusion/vi…

Pedestrian Detection

Learning Local Distortion Visibility From Image Quality Data-sets

2018-03-11 · Navaneeth Kamballur Kottayil, Giuseppe Valenzise, Frederic Dufaux, Irene Cheng

Accurate prediction of local distortion visibility thresholds is critical in many image and video processing applications. Existing methods require an accurate modeling of the human visual system, and are derived through…

Local Distortion

Image-based Visibility Analysis Replacing Line-of-Sight Simulation: An Urban Landmark Perspective

2025-05-17 · Zicheng Fan, Kunihiko Fujiwara, Pengyuan Liu, Fan Zhang 외

Visibility analysis is one of the fundamental analytics methods in urban planning and landscape research, traditionally conducted through computational simulations based on the Line-of-Sight (LoS) principle. However, whe…

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

2026-06-29 · Wentao Feng, Guobei Peng, Wengang Mao, Ryan Wen Liu arxiv

Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monitoring systems. Under hazy conditions, degraded visual information sh…

Image Quality AssessmentImage RestorationObject DetectionImage Dehazing