paper-with-me

홈 › Papers

DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs

2026-04-06 · Dikshant Kukreja, Kshitij Sah, Karan Goyal, Mukesh Mohania, Vikram Goyal arxiv

When asked to describe a molecular diagram, a Vision-Language Model correctly identifies ``a benzene ring with an -OH group.'' When asked to reason about the same image, it answers incorrectly. The model can see but it cannot think about what it sees. We term this the perception-integration gap: a failure where visual information is successfully extracted but lost during downstream reasoning, invisible to single-configuration benchmarks that conflate perception with integration under one accuracy number. To systematically expose such failures, we introduce DISSECT, a 12,000-question diagnostic benchmark spanning Chemistry (7,000) and Biology (5,000). Every question is evaluated under five input modes -- Vision+Text, Text-Only, Vision-Only, Human Oracle, and a novel Model Oracle in which the VLM first verbalizes the image and then reasons from its own description -- yielding diagnostic gaps that decompose performance into language-prior exploitation, visual extraction, perception fidelity, and integration effectiveness. Evaluating 18~VLMs, we find that: (1) Chemistry exhibits substantially lower language-prior exploitability than Biology, confirming molecular visual content as a harder test of genuine visual reasoning; (2) Open-source models consistently score higher when reasoning from their own verbalized descriptions than from raw images, exposing a systematic integration bottleneck; and (3) Closed-source models show no such gap, indicating that bridging perception and integration is the frontier separating open-source from closed-source multimodal capability. The Model Oracle protocol is both model and benchmark agnostic, applicable post-hoc to any VLM evaluation to diagnose integration failures.

📄 PDF Abstract BibTeX arXiv:2604.06250

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

2022-04-23 · Tuomas Oikarinen, Tsui-Wei Weng

In this paper, we propose CLIP-Dissect, a new technique to automatically describe the function of individual hidden neurons inside vision networks. CLIP-Dissect leverages recent advances in multimodal vision/language mod…

Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models

2025-09-25 · Suaiba Amina Salahuddin, Teresa Dorszewski, Marit Almenning Martiniussen, Tone Hovda 외 arxiv

Understanding what deep learning (DL) models learn is essential for the safe deployment of artificial intelligence (AI) in clinical settings. While previous work has focused on pixel-based explainability methods, less at…

DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision

2025-09-23 · Azad Singh, Deepak Mishra arxiv

Self-supervised learning (SSL) has emerged as a powerful paradigm for medical image representation learning, particularly in settings with limited labeled data. However, existing SSL methods often rely on complex archite…

Self-Supervised LearningRepresentation Learning

Describe-and-Dissect: Interpreting Neurons in Vision Networks with Language Models

2024-03-20 · Nicholas Bai, Rahul A. Iyer, Tuomas Oikarinen, Tsui-Wei Weng

In this paper, we propose Describe-and-Dissect (DnD), a novel method to describe the roles of hidden neurons in vision networks. DnD utilizes recent advancements in multimodal deep learning to produce complex natural lan…

Multimodal Deep Learning

Dissected 3D CNNs: Temporal Skip Connections for Efficient Online Video Processing

2020-09-30 · Okan Köpüklü, Stefan Hörmann, Fabian Herzog, Hakan Cevikalp 외

Convolutional Neural Networks with 3D kernels (3D-CNNs) currently achieve state-of-the-art results in video recognition tasks due to their supremacy in extracting spatiotemporal features within video frames. There have b…

Action ClassificationVideo Recognition