paper-with-me

홈 › Papers

Excitation: Momentum For Experts

2026-02-25 · Sagi Shaier arxiv

We propose Excitation, a novel optimization framework designed to accelerate learning in sparse architectures such as Mixture-of-Experts (MoEs). Unlike traditional optimizers that treat all parameters uniformly, Excitation dynamically modulates updates using batch-level expert utilization. It introduces a competitive update dynamic that amplifies updates to highly-utilized experts and can selectively suppress low-utilization ones, effectively sharpening routing specialization. Notably, we identify a phenomenon of "structural confusion" in deep MoEs, where standard optimizers fail to establish functional signal paths; Excitation acts as a specialization catalyst, "rescuing" these models and enabling stable training where baselines remain trapped. Excitation is optimizer-, domain-, and model-agnostic, requires minimal integration effort, and introduces neither additional per-parameter optimizer state nor learnable parameters, making it highly viable for memory-constrained settings. Across language and vision tasks, Excitation consistently improves convergence speed and final performance in MoE models, indicating that active update modulation is a key mechanism for effective conditional computation.

📄 PDF Abstract BibTeX arXiv:2602.21798

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering

2025-03-01 · CVPR 2025 1 · Tianyu Huai, Jie zhou, Xingjiao Wu, Qin Chen 외

Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language tasks (e.g., visual question answering). H…

Continual LearningLanguage ModelingLanguage ModellingLarge Language Model+5

DeepUnifiedMom: Unified Time-series Momentum Portfolio Construction via Multi-Task Learning with Multi-Gate Mixture of Experts

2024-06-13 · Joel Ong, Dorien Herremans

This paper introduces DeepUnifiedMom, a deep learning framework that enhances portfolio management through a multi-task learning approach and a multi-gate mixture of experts. The essence of DeepUnifiedMom lies in its abi…

ManagementMixture-of-ExpertsMulti-Task LearningTime Series

MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts

2024-10-18 · Rachel S. Y. Teo, Tan M. Nguyen

Sparse Mixture of Experts (SMoE) has become the key to unlocking unparalleled scalability in deep learning. SMoE has the potential to exponentially increase parameter count while maintaining the efficiency of the model b…

Language ModelingLanguage ModellingMixture-of-ExpertsObject Recognition

Accelerated Continuous-Time Approximate Dynamic Programming via Data-Assisted Hybrid Control

2022-04-27 · Daniel E. Ochoa, Jorge I. Poveda

We introduce a new closed-loop architecture for the online solution of approximate optimal control problems in the context of continuous-time systems. Specifically, we introduce the first algorithm that incorporates dyna…

Pediatric Wrist Fracture Detection Using Feature Context Excitation Modules in X-ray Images

2024-10-01 · Rui-Yang Ju, Chun-Tse Chien, Enkaer Xieerke, Jen-Shiun Chiang

Children often suffer wrist trauma in daily life, while they usually need radiologists to analyze and interpret X-ray images before surgical treatment by surgeons. The development of deep learning has enabled neural netw…

2D Object DetectionFracture detectionmedical image detectionobject-detection+2