paper-with-me

Papers

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

2026-06-01 · Chuang Ma, Qianying Liu, Tomoyuki Obuchi, Fei Cheng, Wang Yang, Sudong Cai, Shuyuan Zheng, Akiko Aizawa, Sadao Kurohashi arxiv

Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly attended visual information. In this work, we identify a complementary failure mode, spatial lexical bias: adding a spatial relation word to the answer options can attract the model's decision and make the newly added option likely to be selected. Using nine open-weight MLLMs, we show that this phenomenon is widely observed. In particular, models can answer a binary spatial question correctly, yet consistently select an incorrect third spatial option once it is added to the answer set. We isolate such binary-stable but ternary-fragile cases as diagnostic examples and leverage mechanistic interpretability tools, revealing that a substantial part of the failure instead originates on the language side rather than the visual side: visual attention analyses and residual-stream probes show the correct spatial relation remains internally available on these failures, while irrelevant-option controls, activation patching, and sparse component interventions trace the bias to specific LLM-side channels and neurons. Based on this finding, we show that a lightweight LLM-only DPO update on tiny single-object-pair synthetic data mitigates the bias, lifting four-way robust accuracy by up to 100 points on synthetic data, and by 68.0, 32.6, and 20.1 points on broader evaluation datasets WhatsUp, SpatialMQA-Direct, and VSR.

📄 PDF Abstract BibTeX arXiv:2606.01914

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Towards Robustness against Typographic Attack with Training-free Concept Localization

2026-07-02 · Bohan Liu, Wenqian Ye, Guangzhi Xiong, Zhenghao He 외 arxiv

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a cr…

Visual Question AnsweringAutonomous Driving

The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation

2026-05-07 · Hoin Jung, Xiaoqian Wang arxiv

While Multimodal Large Language Models (MLLMs) are increasingly integrated with Retrieval-Augmented Generation (RAG) to mitigate hallucinations, the introduction of external documents can conceal severe failure modes at …

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

2025-10-14 · Noor Islam S. Mohammad, Uluğ Bayazıt arxiv

We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce \textbf{Attribution Graphs} (AGs), which extend GradCAM++ to circuit-level representations, and \textbf{Causal Prob…

Adversarial Robustness

Why do LLaVA Vision-Language Models Reply to Images in English?

2024-07-02 · Musashi Hinck, Carolin Holtermann, Matthew Lyle Olson, Florian Schneider 외

We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significantly increases the likelihood of the model…

Language ModelingLanguage Modelling

Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas

2024-05-09 · Azarakhsh Jalalvand, SangKyeun Kim, Jaemin Seo, Qiming Hu 외

A non-linear system governed by multi-spatial and multi-temporal physics scales cannot be fully understood with a single diagnostic, as each provides only a partial view, leading to information loss. Combining multiple d…

AstronomyDiagnosticImage EnhancementSuper-Resolution