paper-with-me

홈 › Papers

Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity

2026-01-31 · Fanqi Yan, Dung Le, Trang Pham, Huy Nguyen, Nhat Ho arxiv

Contaminated mixture of experts (MoE) is motivated by transfer learning methods where a pre-trained model, acting as a frozen expert, is integrated with an adapter model, functioning as a trainable expert, in order to learn a new task. Despite recent efforts to analyze the convergence behavior of parameter estimation in this model, there are still two unresolved problems in the literature. First, the contaminated MoE model has been studied solely in regression settings, while its theoretical foundation in classification settings remains absent. Second, previous works on MoE models for classification capture pointwise convergence rates for parameter estimation without any guaranty of minimax optimality. In this work, we close these gaps by performing, for the first time, the convergence analysis of a contaminated mixture of multinomial logistic experts with homogeneous and heterogeneous structures, respectively. In each regime, we characterize uniform convergence rates for estimating parameters under challenging settings where ground-truth parameters vary with the sample size. Furthermore, we also establish corresponding minimax lower bounds to ensure that these rates are minimax optimal. Notably, our theories offer an important insight into the design of contaminated MoE, that is, expert heterogeneity yields faster parameter estimation rates and, therefore, is more sample-efficient than expert homogeneity.

📄 PDF Abstract BibTeX arXiv:2602.00939

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts

2024-10-16 · Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian 외

We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prompts, which can be formulated as experts,…

Mixture-of-Expertsparameter estimationPrompt Learning

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

2025-05-24 · Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian 외

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contaminati…

Mixture-of-Experts

Robust W-GAN-Based Estimation Under Wasserstein Contamination

2021-01-20 · Zheng Liu, Po-Ling Loh

Robust estimation is an important problem in statistics which aims at providing a reasonable estimator when the data-generating distribution lies within an appropriately defined ball around an uncontaminated distribution…

regression

Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs

2026-05-19 · Pierre Boudart, Pierre Gaillard, Alessandro Rudi arxiv

We study reinforcement learning for episodic Markov Decision Processes (MDPs) whose transitions are modelled by a multinomial logistic (MNL) model. Existing algorithms for MNL mixture MDPs yield a regret of $\smash{\tild…

Reinforcement Learning

A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts

2023-10-22 · Huy Nguyen, Pedram Akbarian, TrungTin Nguyen, Nhat Ho

Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications. From a theoretical perspective, whil…

Density EstimationMixture-of-Expertsparameter estimationregression