paper-with-me

홈 › Papers

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction

2026-01-16 · Xiaojie Xia, Huigang Zhang, Chaoliang Zhong, Jun Sun, Yusuke Oishi arxiv

Transformer architectures deliver state-of-the-art accuracy via dense full-attention, but their quadratic time and memory complexity with respect to sequence length limits practical deployment. Linear attention mechanisms offer linear or near-linear scaling yet often incur performance degradation. Hybrid models that integrate full and linear attention layers promise a balance between efficiency and expressiveness, but face two major challenges: training such hybrid models from scratch is computationally expensive, and manually designing the optimal placement of attention types is highly nontrivial. We propose DtR (Distill-then-Replace), which first transfers weights from the pretrained full-attention modules to its linear attention counterparts through blockwise local distillation, and then applies a greedy layer replacement strategy that iteratively substitutes full attention blocks with linear ones while monitoring validation performance on the target task. DtR yields a task-specific hybrid model in a single efficient pass, without costly re-training or neural architecture search, and can be applied to any pretrained full-attention backbone for diverse downstream tasks.

📄 PDF Abstract BibTeX arXiv:2601.11667

Code (0)

등록된 구현이 없습니다.

Tasks

Neural Architecture Search

Similar Papers 제목 키워드 기반

Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement

2025-03-20 · CVPR 2025 1 · Shu Yang, Chengting Yu, Lei Liu, Hanzhi Ma 외

Spiking Neural Networks (SNNs) have garnered considerable attention as a potential alternative to Artificial Neural Networks (ANNs). Recent studies have highlighted SNNs' potential on large-scale datasets. For SNN traini…

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

2026-06-23 · Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen 외 arxiv

Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level …

HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation

2026-05-16 · Yihao Liang, Niraj K. Jha arxiv

Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmarks suggest that this works but they hide…

Visual Reasoning

Effective Distillation to Hybrid xLSTM Architectures

2026-03-16 · Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied, Anamaria-Roberta Hartl 외 arxiv

There have been numerous attempts to distill quadratic attention-based large language models (LLMs) into sub-quadratic linearized architectures. However, despite extensive research, such distilled models often fail to ma…

Joint graph entropy knowledge distillation for point cloud classification and robustness against corruptions

2025-09-26 · Zhiqiang Tian, Weigang Li, Junwei Hu, Chunhua Deng arxiv

Classification tasks in 3D point clouds often assume that class events \replaced{are }{follow }independent and identically distributed (IID), although this assumption destroys the correlation between classes. This \repla…

Point Cloud ClassificationKnowledge DistillationPoint Clouds