paper-with-me

홈 › Papers

Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA

2025-12-16 · Esmail Gumaan arxiv

The choice of attention mechanism in Transformer models involves a critical trade-off between modeling quality and inference efficiency. Multi-Head Attention (MHA) offers the best quality but suffers from large Key-Value (KV) cache memory requirements during inference. Multi-Query Attention (MQA) and Grouped-Query Attention (GQA) reduce memory usage but often at the cost of model performance. In this work, we propose Mixture of Attention Schemes (MoAS), a novel architecture that dynamically selects the optimal attention scheme (MHA, GQA, or MQA) for each token via a learned router. We demonstrate that dynamic routing performs better than static averaging of schemes and achieves performance competitive with the MHA baseline while offering potential for conditional compute efficiency. Experimental results on WikiText-2 show that dynamic routing (val loss 2.3074) outperforms a static mixture (2.3093), validating the effectiveness of the proposed method. Our code is available at https://github.com/Esmail-ibraheem/Mixture-of-Attention-Schemes-MoAS.

📄 PDF Abstract BibTeX arXiv:2512.20650

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation

2026-05-18 · Ronyu Zhang, Aosong Cheng, Gaole Dai, Yulin Luo 외 arxiv

Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture-biased backbones risk error accumulation and catastrophic forgettin…

Semantic SegmentationTest-time Adaptation

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

2025-08-28 · Huong Ngo, Matt Deitke, Martijn Bartelds, Sarah Pratt 외 arxiv

Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series o…

Speech Recognition

Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation

2024-05-26 · Rongyu Zhang, Aosong Cheng, Yulin Luo, Gaole Dai 외

Continual Test-Time Adaptation (CTTA), which aims to adapt the pre-trained model to ever-evolving target domains, emerges as an important task for vision models. As current vision models appear to be heavily biased towar…

feature selectionMixture-of-ExpertsTest-time Adaptation

Non-elitist Evolutionary Multi-objective Optimizers Revisited

2020-09-30 · Ryoji Tanabe, Hisao Ishibuchi

Since around 2000, it has been considered that elitist evolutionary multi-objective optimization algorithms (EMOAs) always outperform non-elitist EMOAs. This paper revisits the performance of non-elitist EMOAs for bi-obj…

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling

2025-08-31 · Junfeng Ran, Guangxiang Zhao, Yuhan Wu, Dawei Zhu 외 arxiv

The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models re…