paper-with-me

홈 › Papers

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

2025-11-30 · Cen Lu, Yung-Chen Tang, Andrea Cavallaro arxiv

Large Vision-Language Models (LVLMs) have shown impressive multimodal understanding capabilities, yet the structures that sustain their functionality remain poorly understood from a mechanistic interpretability standpoint. We propose Consistently Activated Neurons (CAN), a progressive neuron ablation method to identify critical neurons whose removal triggers catastrophic collapse, and use it to investigate structural vulnerabilities in representative 7B LVLMs. Experiments reveal that catastrophic collapse can be triggered by ablating as few as four neurons in \texttt{LLaVA-1.5-7b-hf} and a few thousand in \texttt{InstructBLIP-vicuna-7b}, both representing a small fraction of model parameters. Notably, critical neurons are predominantly localized in the language model, particularly in its down-projection layer, rather than in the vision components. We also observe a consistent two-stage collapse pattern: initial expressive degradation followed by sudden, complete collapse. These findings reveal that LVLM functionality depends on a sparse subset of neurons concentrated in the language backbone, offering mechanistic insights into how their functionality is structured and where these models are most vulnerable.

📄 PDF Abstract BibTeX arXiv:2512.00918

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCAN: Sparse Circuit Anchor Interpretable Neuron for Lifelong Knowledge Editing

2026-03-16 · Yuhuan Liu, Haitian Zhong, Xinyuan Xia, Qiang Liu 외 arxiv

Large Language Models (LLMs) often suffer from catastrophic forgetting and collapse during sequential knowledge editing. This vulnerability stems from the prevailing dense editing paradigm, which treats models as black b…

knowledge editing

Emergent Slow Thinking in LLMs as Inverse Tree Freezing

2025-09-28 · Sihan Hu, Xiansheng Cai, Yuan Huang, Zhiyuan Yao 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) enables large language models to acquire slow, multi-step reasoning from sparse final-answer signals. We provide a statistical-physics picture of this emergence. We s…

Reinforcement Learning

Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse

2026-03-18 · Dip Roy, Rajiv Misra, Sanjay Kumar Singh arxiv

Extreme neural network sparsification (90% activation reduction) presents a critical challenge for mechanistic interpretability: understanding whether interpretable features survive aggressive compression. This work inve…

Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Models

2026-06-18 · Hongliang Liu arxiv

Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collaps…

Von Economo neurons enable reliable social skill acquisition in recurrent spiking neural networks: a computational account with clinical predictions

2026-05-17 · Esila Keskin arxiv

Von Economo neurons (VENs) are selectively lost in behavioural-variant frontotemporal dementia (bvFTD) and reduced in autism spectrum conditions (ASC), yet their computational role in social learning remains unexplained.…

Binary Classification