paper-with-me

홈 › Papers

One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers

2025-11-10 · Georgiy Shakirov, Albert Arakelov arxiv

A common practice in heterogeneous graph neural networks (HGNNs) is to condition parameters on node/edge types, assuming types reflect semantic roles. However, this can cause overreliance on surface-level labels and impede cross-type knowledge transfer. We explore integrating Mixture-of-Experts (MoE) into HGNNs--a direction underexplored despite MoE's success in homogeneous settings. Crucially, we question the need for type-specific experts. We propose Homogeneous Expert Routing (HER), an MoE layer for Heterogeneous Graph Transformers (HGT) that stochastically masks type embeddings during routing to encourage type-agnostic specialization. Evaluated on IMDB, ACM, and DBLP for link prediction, HER consistently outperforms standard HGT and a type-separated MoE baseline. Analysis on IMDB shows HER experts specialize by semantic patterns (e.g., movie genres) rather than node types, confirming routing is driven by latent semantics. Our work demonstrates that regularizing type dependence in expert routing yields more generalizable, efficient, and interpretable representations--a new design principle for heterogeneous graph learning.

📄 PDF Abstract BibTeX arXiv:2511.07603

Code (0)

등록된 구현이 없습니다.

Tasks

Link PredictionGraph Learning

Similar Papers 제목 키워드 기반

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

2026-05-12 · Sagi Ahrac, Noya Hochwald, Mor Geva arxiv

Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse onto few experts and auxiliary load-balancing losses can reduce specializ…

HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts

2023-12-12 · Giang Do, Khiem Le, Quang Pham, TrungTin Nguyen 외

By routing input tokens to only a few split experts, Sparse Mixture-of-Experts has enabled efficient training of large language models. Recent findings suggest that fixing the routers can achieve competitive performance …

Mixture-of-Experts

Self-Routing: Parameter-Free Expert Routing from Hidden States

2026-04-01 · Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli arxiv

Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to expert assignments. In this work, we ask whet…

IR3DE: A Linear Router for Large Language Models

2026-06-04 · Eros Fanì, Oğuzhan Ersoy arxiv

Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialized tasks via domain-expert LLMs. With the ever-growing list of availa…

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

2026-02-22 · Yuqi Xu, Rizhen Hu, Zihan Liu, Mou Sun 외 arxiv

Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searching for an optimal routing policy within …