paper-with-me

홈 › Papers

Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts

2026-01-29 · Yingfa Chen, Zhen Leng Thai, Zihan Zhou, Zhu Zhang, Xingyu Shen, Shuo Wang, Chaojun Xiao, Xu Han, Zhiyuan Liu arxiv

Hybrid Transformer architectures, which combine softmax attention blocks and recurrent neural networks (RNNs), have shown a desirable performance-throughput tradeoff for long-context modeling, but their adoption and studies are hindered by the prohibitive cost of large-scale pre-training from scratch. Some recent studies have shown that pre-trained softmax attention blocks can be converted into RNN blocks through parameter transfer and knowledge distillation. However, these transfer methods require substantial amounts of training data (more than 10B tokens), and the resulting hybrid models also exhibit poor long-context performance, which is the scenario where hybrid models enjoy significant inference speedups over Transformer-based models. In this paper, we present HALO (Hybrid Attention via Layer Optimization), a pipeline for distilling Transformer models into RNN-attention hybrid models. We then present HypeNet, a hybrid architecture with superior length generalization enabled by a novel position encoding scheme (named HyPE) and various architectural modifications. We convert the Qwen3 series into HypeNet using HALO, achieving performance comparable to the original Transformer models while enjoying superior long-context performance and efficiency. The conversion requires just 2.3B tokens, less than 0.01% of their pre-training data

📄 PDF Abstract BibTeX arXiv:2601.22156

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation

2026-04-11 · Ruibin Li, Tao Yang, Fangzhou Ai, Tianhe Wu 외 arxiv

Streaming video generation (SVG) distills a pretrained bidirectional video diffusion model into an autoregressive model equipped with sliding window attention (SWA). However, SWA inevitably loses distant history during l…

Computational EfficiencyModel CompressionVideo Generation

ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers

2026-01-07 · Mohsen Ghafoorian, Amirhossein Habibian arxiv

Recent advances in video diffusion models have shifted towards transformer-based architectures, achieving state-of-the-art video generation but at the cost of quadratic attention complexity, which severely limits scalabi…

Video Generation

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction

2026-01-16 · Xiaojie Xia, Huigang Zhang, Chaoliang Zhong, Jun Sun 외 arxiv

Transformer architectures deliver state-of-the-art accuracy via dense full-attention, but their quadratic time and memory complexity with respect to sequence length limits practical deployment. Linear attention mechanism…

Neural Architecture Search

Explainable RL Policies by Distilling to Locally-Specialized Linear Policies with Voronoi State Partitioning

2025-11-17 · Senne Deproost, Dennis Steckelmacher, Ann Nowé arxiv

Deep Reinforcement Learning is one of the state-of-the-art methods for producing near-optimal system controllers. However, deep RL algorithms train a deep neural network, that lacks transparency, which poses challenges w…

Reinforcement Learning

Morphing into Hybrid Attention Models

2026-06-29 · Disen Lan, Jianbin Zheng, Yuxi Ren, Xin Xia 외 hf

Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the effectiveness of Transformer-to-hybrid co…