paper-with-me

홈 › Papers

Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models

2026-03-28 · Junhyeok Lee, Kyu Sung Choi arxiv

Mixture-of-Experts (MoE) language models are universally sensitive to demographic content at the routing level, yet exploiting this sensitivity for fairness control is structurally limited. We introduce Fairness-Aware Routing Equilibrium (FARE), a diagnostic framework designed to probe the limits of routing-level stereotype intervention across diverse MoE architectures. FARE reveals that routing-level preference shifts are either unachievable (Mixtral, Qwen1.5, Qwen3), statistically non-robust (DeepSeekMoE), or accompanied by substantial utility cost (OLMoE, -4.4%p CrowS-Pairs at -6.3%p TQA). Critically, even where log-likelihood preference shifts are robust, they do not transfer to decoded generation: expanded evaluations on both non-null models yield null results across all generation metrics. Group-level expert masking reveals why: bias and core knowledge are deeply entangled within expert groups. These findings indicate that routing sensitivity is necessary but insufficient for stereotype control, and identify specific architectural conditions that can inform the design of more controllable future MoE systems.

📄 PDF Abstract BibTeX arXiv:2603.27141

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

2026-07-28 · Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan 외 arxiv

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visua…

Image Reconstruction

Probing Routing-Conditional Calibration in Attention-Residual Transformers

2026-05-11 · Wenhao Liang, Lin Yue, Wei Emma Zhang, Miao Xu 외 arxiv

Post-hoc calibration is usually evaluated as a function of logits or softmax confidence alone, even as routing-augmented architectures increasingly accompany predictions with sample-specific internal routing traces and p…

Edge-aware Decoding for Neural Asymmetric Routing

2026-06-01 · Li Liang, Jinbiao Chen, Zizhen Zhang arxiv

Neural asymmetric routing models increasingly encode directionality through matrix representations and asymmetry-aware attention. The final routing action, however, is not a node in isolation but a directed transition ch…

Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing

2026-06-28 · Yoosung Hong arxiv

Role-semantic assignments provide priors over how heterogeneous agents may coordinate, but cooperative MARL systems instead settle on conventions through decentralized, non-stationary learning, with no guarantee that the…

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

2025-11-18 · Ankush Kadu, Aswanth Krishnan arxiv

We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commit to a wrong approach early and exhaust the step budget, the post-failure traj…