paper-with-me

홈 › Papers

Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making

2025-10-19 · Larkin Liu, Jalal Etesami arxiv

We explore the use of expert-guided bandit learning, which we refer to as online mixture-of-experts (OMoE). In this setting, given a context, a candidate committee of experts must determine how to aggregate their outputs to achieve optimal results in terms of aggregate accuracy. We propose two algorithms to address this problem. The first algorithm combines aggregate voting with UCB-driven successive elimination, efficiently pruning suboptimal exploration actions. The second algorithm employs an online weighted-majority-voting mechanism, leveraging the respective voting power of each expert proportional to their predictive power. We derive theoretical guarantees for the regret properties in the bandit setting under ideal circumstances, and empirical results are provided accordingly. As a modern study on applications, these methods are applied to the online fine-tuning of a set of expert large language models (LLMs), where after each response, the generative LLM dynamically reweighs its set of experts and/or selects the optimal committee of experts to generate the most accurate response. Our results introduce new methodologies and no-regret guarantees for combining multiple experts to improve on the performance of the an aggregate model overall.

📄 PDF Abstract BibTeX arXiv:2510.21788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Replicable Online Learning

2024-11-20 · Saba Ahmadi, Siddharth Bhandari, Avrim Blum

We investigate the concept of algorithmic replicability introduced by Impagliazzo et al. 2022, Ghazi et al. 2021, Ahn et al. 2024 in an online setting. In our model, the input sequence received by the online learner is g…

Mixture of Online and Offline Experts for Non-stationary Time Series

2022-02-12 · Zhilin Zhao, Longbing Cao, Yuanyu Wan

We consider a general and realistic scenario involving non-stationary time series, consisting of several offline intervals with different distributions within a fixed offline time horizon, and an online interval that con…

Time SeriesTransfer Learning

Near-Optimal Algorithms for Private Online Optimization in the Realizable Regime

2023-02-27 · Hilal Asi, Vitaly Feldman, Tomer Koren, Kunal Talwar

We consider online learning problems in the realizable setting, where there is a zero-loss solution, and propose new Differentially Private (DP) algorithms that obtain near-optimal regret bounds. For the problem of onlin…

Continuous Prediction with Experts' Advice

2022-06-01 · Victor Sanches Portella, Christopher Liaw, Nicholas J. A. Harvey

Prediction with experts' advice is one of the most fundamental problems in online learning and captures many of its technical challenges. A recent line of work has looked at online learning through the lens of differenti…

Prediction

Efficient and Optimal Fixed-Time Regret with Two Experts

2022-03-15 · Laura Greenstreet, Nicholas J. A. Harvey, Victor Sanches Portella

Prediction with expert advice is a foundational problem in online learning. In instances with $T$ rounds and $n$ experts, the classical Multiplicative Weights Update method suffers at most $\sqrt{(T/2)\ln n}$ regret when…

Vocal Bursts Valence Prediction