paper-with-me

홈 › Papers

Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs

2025-10-13 · João Paulo Cardoso de Lima, Marc Dietrich, Jeronimo Castrillon, Asif Ali Khan arxiv

Structured sparsity enables deploying large language models (LLMs) on resource-constrained systems. Approaches like dense-to-sparse fine-tuning are particularly compelling, achieving remarkable structured sparsity by reducing the model size by over 6.7x, while still maintaining acceptable accuracy. Despite this reduction, LLM inference, especially the decode stage being inherently memory-bound, is extremely expensive on conventional Von-Neumann architectures. Compute-in-memory (CIM) architectures mitigate this by performing computations directly in memory, and when paired with sparse LLMs, enable storing and computing the entire model in memory, eliminating the data movement on the off-chip bus and improving efficiency. Nonetheless, naively mapping sparse matrices onto CIM arrays leads to poor array utilization and diminished computational efficiency. In this paper, we present an automated framework with novel mapping and scheduling strategies to accelerate sparse LLM inference on CIM accelerators. By exploiting block-diagonal sparsity, our approach improves CIM array utilization by over 50%, achieving more than 4x reduction in both memory footprint and the number of required floating-point operations.

📄 PDF Abstract BibTeX arXiv:2510.11192

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization

2025-10-07 · Lawrence Liu, Alexander Liu, Mengdi Wang, Tuo Zhao 외 arxiv

Large language models (LLMs) present significant deployment challenges due to their immense computational and memory requirements. While semi-structured pruning, particularly 2:4 sparsity, offers a path to practical hard…

Model Compression

Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

2026-05-19 · Wenhu Zhang, Yiming Wu, Huanyu Wang, Yaoyang Liu 외 arxiv

Diffusion Language Models (DLMs) enable globally coherent, bidirectional, and controllable text generation, offering advantages over traditional autoregressive LLMs, while scaling to ultra-long sequences remains costly. …

Video GenerationText Generation

Hierarchical Sparse Plus Low Rank Compression of LLM

2025-12-19 · Pawan Kumar, Aditi Gupta arxiv

Modern large language models (LLMs) place extraordinary pressure on memory and compute budgets, making principled compression indispensable for both deployment and continued training. We present Hierarchical Sparse Plus …

XAttention: Block Sparse Attention with Antidiagonal Scoring

2025-03-20 · Ruyi Xu, Guangxuan Xiao, Haofeng Huang, Junxian Guo 외

Long-Context Transformer Models (LCTMs) are vital for real-world applications but suffer high computational costs due to attention's quadratic complexity. Block-sparse attention mitigates this by focusing computation on …

Video GenerationVideo Understanding

A Decomposition Framework for Certifiably Optimal Orthogonal Sparse PCA

2026-03-01 · Difei Cheng, Qiao Hu arxiv

Sparse Principal Component Analysis (SPCA) is an important technique for high-dimensional data analysis, improving interpretability by imposing sparsity on principal components. However, existing methods often fail to si…