paper-with-me

홈 › Papers

Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models

2024-04-07 · Jordan Dotzel, Yash Akhauri, Ahmed S. AbouElhamayed, Carly Jiang, Mohamed Abdelfattah, Zhiru Zhang

Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce compute on an input-by-input basis. These methods improve over static methods by exploiting the variance across individual inputs, which has steadily grown with the exponential increase in training data. Yet, the increasing depth within modern models, currently with hundreds of layers, has opened opportunities for dynamic layer sparsity, which skips the computation for entire layers. In this work, we explore the practicality of layer sparsity by profiling residual connections and establish the relationship between model depth and layer sparsity. For example, the residual blocks in the OPT-66B model have a median contribution of 5% to its output. We then take advantage of this dynamic sparsity and propose Radial Networks, which perform token-level routing between layers guided by a trained router module. These networks can be used in a post-training distillation from sequential networks or trained from scratch to co-learn the router and layer weights. They enable scaling to larger model sizes by decoupling the number of layers from the dynamic depth of the network, and their design allows for layer reuse. By varying the compute token by token, they reduce the overall resources needed for generating entire sequences. Overall, this leads to larger capacity networks with significantly lower compute and serving costs for large language models.

📄 PDF Abstract BibTeX arXiv:2404.04900

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing

2025-06-04 · Ruihan Jin, Pengpeng Shao, Zhengqi Wen, Jinyang Wu 외

The rapid advancements in large language models (LLMs) have led to the emergence of routing techniques, which aim to efficiently select the optimal LLM from diverse candidates to tackle specific tasks, optimizing perform…

MACRO: Markov Chain Routing of Transformer Layers

2026-08-06 · Paweł Batorski, Abtin Pourhadi, Akylgali Aitaza, Przemysław Spurek 외 arxiv

Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers involving layer repetitions, skips and other moves, can improve performa…

Adaptive Routing Between Capsules

2019-11-19 · Qiang Ren, Shaohua Shang, Lianghua He

Capsule network is the most recent exciting advancement in the deep learning field and represents positional information by stacking features into vectors. The dynamic routing algorithm is used in the capsule network, ho…

OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning

2024-03-20 · CVPR 2024 1 · Xinyu Geng, JiaMing Wang, Jiawei Gong, Yuerong Xue 외

Redundancy is a persistent challenge in Capsule Networks (CapsNet),leading to high computational costs and parameter counts. Although previous works have introduced pruning after the initial capsule layer, dynamic routin…

XnODR and XnIDR: Two Accurate and Fast Fully Connected Layers For Convolutional Neural Networks

2021-11-21 · Jian Sun, Ali Pourramezan Fard, Mohammad H. Mahoor

Capsule Network is powerful at defining the positional relationship between features in deep neural networks for visual recognition tasks, but it is computationally expensive and not suitable for running on mobile device…

BinarizationImage Classification