paper-with-me

홈 › Papers

SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks

2024-02-14 · Jiwon Song, Kyungseok Oh, Taesu Kim, HyungJun Kim, Yulhwa Kim, Jae-Joon Kim

Large language models (LLMs) have proven to be highly effective across various natural language processing tasks. However, their large number of parameters poses significant challenges for practical deployment. Pruning, a technique aimed at reducing the size and complexity of LLMs, offers a potential solution by removing redundant components from the network. Despite the promise of pruning, existing methods often struggle to achieve substantial end-to-end LLM inference speedup. In this paper, we introduce SLEB, a novel approach designed to streamline LLMs by eliminating redundant transformer blocks. We choose the transformer block as the fundamental unit for pruning, because LLMs exhibit block-level redundancy with high similarity between the outputs of neighboring blocks. This choice allows us to effectively enhance the processing speed of LLMs. Our experimental results demonstrate that SLEB outperforms previous LLM pruning methods in accelerating LLM inference while also maintaining superior perplexity and accuracy, making SLEB as a promising technique for enhancing the efficiency of LLMs. The code is available at: https://github.com/jiwonsong-dev/SLEB.

📄 PDF Abstract BibTeX arXiv:2402.09025

Code (1)

jiwonsong-dev/sleb 공식 구현

Methods 이 논문이 사용한 방법론

Pruning 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

AIvril: AI-Driven RTL Generation With Verification In-The-Loop

2024-09-03 · Mubashir ul Islam, Humza Sami, Pierre-Emmanuel Gaillardon, Valerio Tenace

Large Language Models (LLMs) are computational models capable of performing complex natural language processing tasks. Leveraging these capabilities, LLMs hold the potential to transform the entire hardware design stack,…

Code Generation

ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models

2025-09-20 · Haoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li 외 arxiv

Large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks. Understanding how LLMs internally represent knowledge remains a significant challenge. Despite Sparse Autoe…

Accelerate Speculative Decoding with Sparse Computation in Verification

2025-12-26 · Jikai Wang, Jianchao Tan, Yuxuan Hu, Jiayu Qin 외 arxiv

Speculative decoding accelerates autoregressive language model inference by verifying multiple draft tokens in parallel. However, the verification stage often becomes the dominant computational bottleneck, especially for…

Mathematical ReasoningQuestion Answering

MCCoder: Streamlining Motion Control with LLM-Assisted Code Generation and Rigorous Verification

2024-10-19 · Yin Li, Liangwei Wang, Shiyuan Piao, Boo-Ho Yang 외

Large Language Models (LLMs) have shown considerable promise in code generation. However, the automation sector, especially in motion control, continues to rely heavily on manual programming due to the complexity of task…

Code GenerationRAGRetrieval-augmented Generation

BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking

2025-02-22 · Yuxuan Liu, Hongda Sun, Wenya Guo, Xinyan Xiao 외

Complex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex …

Fact Checking