paper-with-me

Papers

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

2026-06-30 · Chih-Ting Liao, Fei Shen, Xin Cao, Tat-Seng Chua arxiv

The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-language model (VLM) actually grounds in the image. We show this on spatial reasoning, where the error is invisible to both probing and steering yet exposed by a one-line causal control: replacing the image with a gray blank. Probes decode the within-axis answer at 73--97% across axes, and a training-free projection lifts a near-chance axis from 59% to 79%, exactly the signature of unlocking latent knowledge. The blank-image arbiter refutes it, revealing three grounding regimes that probing conflates: an axis can be grounded (vision-dependent, correct), a prior (vision-independent, with its decode and its apparent recovery a directional default rather than perception), or, surprisingly, inverted: decodable, causally controllable, but deployed with the wrong sign, so the model scores below chance and the error requires looking. The taxonomy holds across the studied VLMs: in fourteen models spanning six language-model families and 2B--27B, horizontal is grounded, vertical is a prior, and depth is inverted, with the inversion emerging at scale within families. The decode-versus-deploy inversion replicates on seven of eight models across five families, and the minimal edit that re-deploys it varies with geometry: a training-free rotation matches a trained edit on the cleanest model, while distributed inversions need a trained low-rank edit, tracing a per-model correction-complexity spectrum. The cheap, self-calibrating arbiter cleanly separates grounded perception, inverted perception, and prior substitution; we argue it should be a default control for latent-knowledge and steering claims in VLMs.

📄 PDF Abstract BibTeX arXiv:2606.31257

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

2026-08-24 · Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li 외 arxiv

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …

Object LocalizationSpatial Reasoning

DamageArbiter: A Multimodal Arbitration Framework for Disaster Damage Assessment from Street-View Imagery

2026-03-16 · Yifan Yang, Lei Zou, Wenjing Gong, Kani Fu 외 arxiv

Analyzing street-view imagery with computer vision models offers a promising approach for rapid, hyperlocal disaster damage assessment, but existing approaches typically rely on black-box pre-trained vision models, which…

WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents

2026-01-29 · Yao Zhang, Shijie Tang, Zeyu Li, Zhen Han 외 arxiv

Web agents hold great potential for automating complex computer tasks, yet their interactions involve long-horizon, sequential decision-making with irreversible actions. In such settings, outcome-based supervision is spa…

Reinforcement LearningText Generation

From Pixels to Hierarchical Sequences: Quadtree Mask Encoding for Vision-Language Binary Change Detection

2026-09-09 · Xiao An, Ruikang Zhang, Chen Zhong, Xuli Shen 외 arxiv

Dense change detection in remote sensing requires vision-language models (VLMs) to compare bi-temporal images and generate accurate pixel-level masks. Existing VLMs are largely confined to change captioning outputs, and …

Change Detection

A Risk-Neutral Neural Operator for Arbitrage-Free SPX-VIX Term Structures

2025-11-09 · Jian'an Zhang arxiv

We propose ARBITER, a risk-neutral neural operator for learning joint SPX-VIX term structures under no-arbitrage constraints. ARBITER maps market states to an operator that outputs implied volatility and variance curves …