paper-with-me

홈 › Papers

Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts

2024-10-16 · Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian, Nhat Ho

We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prompts, which can be formulated as experts, to fine-tune a large-scaled pre-trained model for learning downstream tasks. There are two fundamental challenges emerging from the analysis: (i) the proportion in the mixture of the pre-trained model and the prompt may converge to zero where the prompt vanishes during the training; (ii) the algebraic interaction among parameters of the pre-trained model and the prompt can occur via some partial differential equation and decelerate the prompt learning. In response, we introduce a distinguishability condition to control the previous parameter interaction. Additionally, we also consider various types of expert structures to understand their effects on the parameter estimation. In each scenario, we provide comprehensive convergence rates of parameter estimation along with the corresponding minimax lower bounds.

📄 PDF Abstract BibTeX arXiv:2410.12258

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-Expertsparameter estimationPrompt Learning

Similar Papers 제목 키워드 기반

Improving Minimax Estimation Rates for Contaminated Mixture of Multinomial Logistic Experts via Expert Heterogeneity

2026-01-31 · Fanqi Yan, Dung Le, Trang Pham, Huy Nguyen 외 arxiv

Contaminated mixture of experts (MoE) is motivated by transfer learning methods where a pre-trained model, acting as a frozen expert, is integrated with an adapter model, functioning as a trainable expert, in order to le…

Transfer Learning

Singularity structures and impacts on parameter estimation in finite mixtures of distributions

2016-09-09 · Nhat Ho, XuanLong Nguyen

Singularities of a statistical model are the elements of the model's parameter space which make the corresponding Fisher information matrix degenerate. These are the points for which estimation techniques such as the max…

parameter estimation

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

2025-05-24 · Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian 외

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contaminati…

Mixture-of-Experts

Convergence Rates for Gaussian Mixtures of Experts

2019-07-09 · Nhat Ho, Chiao-Yu Yang, Michael. I. Jordan

We provide a theoretical treatment of over-specified Gaussian mixtures of experts with covariate-free gating networks. We establish the convergence rates of the maximum likelihood estimation (MLE) for these models. Our p…

parameter estimation

Optimal Estimation and Completion of Matrices with Biclustering Structures

2015-12-01 · Chao Gao, Yu Lu, Zongming Ma, Harrison H. Zhou

Biclustering structures in data matrices were first formalized in a seminal paper by John Hartigan (1972) where one seeks to cluster cases and variables simultaneously. Such structures are also prevalent in block modelin…