paper-with-me

홈 › Papers

BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers

2025-10-27 · Ramshankar Bhuvaneswaran, Handan Liu arxiv

The pursuit of efficient Large Language Models (LLMs) has led to increasingly complex techniques like extreme quantization and dynamic routing. While individual benefits of these methods are well-documented, their compositional effects remain poorly understood. This paper introduces BitSkip, a hybrid architectural framework for systematically exploring these interactions. Counter-intuitively, our findings reveal that a simple 8-bit quantized model without Hadamard transform (BitSkip-V1) not only outperforms its more complex 4-bit and Hadamard-enhanced counterparts but also competes the full-precision baseline in quality (perplexity of 1.13 vs 1.19) . The introduction of Hadamard transforms, even at 8-bit precision, catastrophically degraded performance by over 37,000%, tracing fundamental training instability. Our BitSkip-V1 recipe demonstrates superior early-exit characteristics, with layer 18 providing optimal 32.5% speed gain for minimal 4% quality loss.

📄 PDF Abstract BibTeX arXiv:2510.23766

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Predicting Probabilities of Error to Combine Quantization and Early Exiting: QuEE

2024-06-20 · Florence Regol, Joud Chataoui, Bertrand Charpentier, Mark Coates 외

Machine learning models can solve complex tasks but often require significant computational resources during inference. This has led to the development of various post-training computation reduction methods that tackle t…

Quantization

On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance

2024-03-25 · Jaskirat Singh, Emad Fallahzadeh, Bram Adams, Ahmed E. Hassan

Deciding what combination of operators to use across the Edge AI tiers to achieve specific latency and model performance requirements is an open question for MLOps engineers. This study aims to empirically assess the acc…

CPUQuantization

Amortized-Precision Quantization for Early-Exit Vision Transformers

2026-05-08 · Rui Fang, Hsi-Wen Chen, Ming-Syan Chen arxiv

Vision Transformers (ViTs) achieve strong performance across vision tasks, yet their deployment with low-precision early exiting remains fragile. Existing quantization methods assume static full-depth execution, making t…

McQueen : Mixed Precision Quantization of Early Exit Networks

2023-11-20 · British Machine Vision Conference (BMVC) 2023 11 · Utkarsh Saxena; Kaushik Roy

Mixed precision quantization offers a promising way of obtaining the optimal tradeoff between model complexity and accuracy. However, most quantization techniques do not support input adaptive execution of neural network…

Quantization

Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators

2025-12-04 · Alaa Zniber, Arne Symons, Ouassim Karrakchou, Marian Verhelst 외 arxiv

Deployment of dynamic neural networks on edge accelerators requires careful consideration of hardware constraints beyond conventional complexity metrics such as Multiply-Accumulate operations. In Early-Exiting Neural Net…

Neural Architecture Search