paper-with-me

홈 › Papers

VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking

2025-11-24 · Kichang Yang, Seonjun Kim, Minjae Kim, Nairan Zhang, Chi Zhang, Youngki Lee arxiv

Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric, selecting neurons solely by activation magnitude and neglecting how access patterns influence flash performance. We present Neuron Chunking, an I/O-efficient sparsification strategy that operates on chunks (i.e., groups of contiguous neurons in memory) and couples neuron importance with storage access cost. The method models I/O latency through a lightweight abstraction of access contiguity and selects chunks with high utility, defined as neuron importance normalized by estimated latency. By aligning sparsification decisions with the underlying storage behavior, Neuron Chunking improves I/O efficiency by up to 4.65x and 5.76x on Jetson Orin Nano and Jetson AGX Orin, respectively.

📄 PDF Abstract BibTeX arXiv:2511.18692

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EventFlash: Towards Efficient MLLMs for Event-Based Vision

2026-02-03 · Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 외 arxiv

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense…

Event-based vision

Structured Sparsification of Gated Recurrent Neural Networks

2019-11-13 · Ekaterina Lobacheva, Nadezhda Chirkova, Alexander Markovich, Dmitry Vetrov

Recently, a lot of techniques were developed to sparsify the weights of neural networks and to remove networks' structure units, e.g. neurons. We adjust the existing sparsification approaches to the gated recurrent archi…

Language ModelingLanguage Modellingtext-classificationText Classification

FLASHE: Additively Symmetric Homomorphic Encryption for Cross-Silo Federated Learning

2021-09-02 · Zhifeng Jiang, Wei Wang, Yang Liu

Homomorphic encryption (HE) is a promising privacy-preserving technique for cross-silo federated learning (FL), where organizations perform collaborative model training on decentralized data. Despite the strong privacy g…

Federated LearningPrivacy Preserving

Utilizing dynamic sparsity on pretrained DETR

2025-10-10 · Reza Sedghi, Anand Subramoney, David Kappel arxiv

Efficient inference with transformer-based models remains a challenge, especially in vision tasks like object detection. We analyze the inherent sparsity in the MLP layers of DETR and introduce two methods to exploit it …

Object Detection

Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse

2026-03-18 · Dip Roy, Rajiv Misra, Sanjay Kumar Singh arxiv

Extreme neural network sparsification (90% activation reduction) presents a critical challenge for mechanistic interpretability: understanding whether interpretable features survive aggressive compression. This work inve…