paper-with-me

Papers

What Gets Activated: Uncovering Domain and Driver Experts in MoE Language Models

2026-01-15 · Guimin Hu, Meng Li, Qiwei Peng, Lijie Hu, Boyan Xu, Ruichu Cai arxiv

Most interpretability work focuses on layer- or neuron-level mechanisms in Transformers, leaving expert-level behavior in MoE LLMs underexplored. Motivated by functional specialization in the human brain, we analyze expert activation by distinguishing domain and driver experts. In this work, we study expert activation in MoE models across three public domains and address two key questions: (1) which experts are activated, and whether certain expert types exhibit consistent activation patterns; and (2) how tokens are associated with and trigger the activation of specific experts. To answer these questions, we introduce entropy-based and causal-effect metrics to assess whether an expert is strongly favored for a particular domain, and how strongly expert activation contributes causally to the model's output, thus identify domain and driver experts, respectively. Furthermore, we explore how individual tokens are associated with the activation of specific experts. Our analysis reveals that (1) Among the activated experts, some show clear domain preferences, while others exert strong causal influence on model performance, underscoring their decisive roles. (2) tokens occurring earlier in a sentence are more likely to trigger the driver experts, and (3) adjusting the weights of domain and driver experts leads to significant performance gains across all three models and domains. These findings shed light on the internal mechanisms of MoE models and enhance their interpretability.

📄 PDF Abstract BibTeX arXiv:2601.10159

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LAYER: A Quantitative Explainable AI Framework for Decoding Tissue-Layer Drivers of Myofascial Low Back Pain

2025-11-25 · Zixue Zeng, Anthony M. Perti, Tong Yu, Grant Kokenberger 외 arxiv

Myofascial pain (MP) is a leading cause of chronic low back pain, yet its tissue-level drivers remain poorly defined and lack reliable image biomarkers. Existing studies focus predominantly on muscle while neglecting fas…

FareShare: A Tool for Labor Organizers to Estimate Lost Wages and Contest Arbitrary AI and Algorithmic Deactivations

2025-05-13 · Varun Nagaraj Rao, Samantha Dalal, Andrew Schwartz, Amna Liaqat 외

What happens when a rideshare driver is suddenly locked out of the platform connecting them to riders, wages, and daily work? Deactivation-the abrupt removal of gig workers' platform access-typically occurs through arbit…

Leveraging Domain Adaptation for Low-Resource Geospatial Machine Learning

2021-07-11 · Jack Lynch, Sam Wookey

Machine learning in remote sensing has matured alongside a proliferation in availability and resolution of geospatial imagery, but its utility is bottlenecked by the need for labeled data. What's more, many labeled geosp…

BIG-bench Machine LearningDomain Adaptation

Computational methods for cancer driver discovery: A survey

2020-07-02 · Vu Viet Hoang Pham, Lin Liu, Cameron Bracken, Gregory Goodall 외

Motivation: Uncovering the genomic causes of cancer, known as cancer driver genes, is a fundamental task in biomedical research. Cancer driver genes drive the development and progression of cancer, thus identifying cance…

Driver IdentificationSurvey

Driver Fatigue Prediction using Randomly Activated Neural Networks for Smart Ridesharing Platforms

2024-04-16 · Sree Pooja Akula, Mukund Telukunta, Venkata Sriram Siddhardh Nadendla

Drivers in ridesharing platforms exhibit cognitive atrophy and fatigue as they accept ride offers along the day, which can have a significant impact on the overall efficiency of the ridesharing platform. In contrast to t…