paper-with-me

홈 › Papers

Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning

2026-05-31 · Garvin Guo, Yu Chen, Xiang Wang, Shuai Li, Xinpei Zhao, Huaxing Liu, Shuai Dong arxiv

Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the tokens encoding visual evidence; recent analyses, however, reveal a paradox: the tokens are loosely tied to the image and contribute little to the answer. Critically, these analyses treat latent tokens as a single unit, obscuring the source of the gains. We therefore decompose latent tokens into three testable components: latent slots, boundary markers, and format, and develop a state-of-the-art method as a probe under favorable conditions. Across six method-stage settings and four perception-heavy benchmarks, latent slots fail every prediction of the visual-memory account. Strikingly, retaining only the boundary markers preserves 78 to 100% of the gain in several settings, while the model attends to the image more narrowly at latent positions than at answer positions. These results do not support slot contents as recoverable visual memory; much of the benefit is instead associated with marker and format control, together with visual-attention routing. At matched accuracy, methods can still rely on markedly different mechanisms shaped by training supervision. Latent visual reasoning thus needs evaluation not only by accuracy but by what the model actually relies on.

📄 PDF Abstract BibTeX arXiv:2606.01287

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks

2025-12-17 · Vegard Flovik arxiv

Deep neural networks achieve impressive performance but remain difficult to interpret and control. We present SALVE (Sparse Autoencoder-Latent Vector Editing), a unified "discover, validate, and control" framework that b…

Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization

2026-05-07 · Andy Zeyi Liu, Elliot Paquette, John Sous arxiv

Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and operational diagnostics. Using a control…

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

2026-09-03 · Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu 외 hf

Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically com…

LOLAMEME: Logic, Language, Memory, Mechanistic Framework

2024-05-31 · Jay Desai, Xiaobo Guo, Srinivasan H. Sengamedu

The performance of Large Language Models has achieved superhuman breadth with unprecedented depth. At the same time, the language models are mostly black box models and the underlying mechanisms for performance have been…

Language ModelingLanguage Modelling

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

2026-05-29 · Yunpeng Zhou arxiv

Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state evolution in low-capacity regimes remain underexplored. We study fa…

Visual Question AnsweringVisual Reasoning