paper-with-me

홈 › Papers

Online Learning-to-Defer with Varying Experts

2026-05-12 · Dang Hoang Duy, Yannis Montreuil, Maxime Meyer, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi arxiv

Learning-to-Defer (L2D) methods route each query either to a predictive model or to external experts. Real-world deployments require handling streaming data, changing expert availability, shifting expert reliability, and feedback observed only for the selected action. We introduce an online multiclass L2D algorithm that combines queried-action bandit feedback with a dynamically varying pool of experts. Let $N=n+n_e$, let $B$ bound the Frobenius norm of the linear score matrix, and let $ρ$ bound the augmented input norm. Assuming linear calibration and zero surrogate minimizability gap for the projected comparator class, our method achieves expected true-deferral regret $O((BN^{3/2}ρ+1)T^{2/3})$, improving to $O(BN^{3/2}ρ\sqrt T+B^2N^3ρ^2)$ under a concentrated-score condition. The analysis combines an online $\mathcal H$-consistency transfer bound with projected online convex optimization. Experiments on synthetic and real-world datasets demonstrate selective routing under varying expert availability and reliability.

📄 PDF Abstract BibTeX arXiv:2605.12340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Why Ask One When You Can Ask $k$? Two-Stage Learning-to-Defer to the Top-$k$ Experts

2025-04-17 · Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

Although existing Learning-to-Defer (L2D) frameworks support multiple experts, they allocate each query to a single expert, limiting their ability to leverage collective expertise in complex decision-making scenarios. To…

Decision Making

Budgeted Multiple-Expert Deferral

2025-10-30 · Giulia DeSalvo, Clara Mohri, Mehryar Mohri, Yutao Zhong arxiv

Learning to defer uncertain predictions to costly experts offers a powerful strategy for improving the accuracy and efficiency of machine learning systems. However, standard training procedures for deferral algorithms ty…

Active Learning

Expert-Agnostic Learning to Defer

2025-02-14 · Joshua Strong, Pramit Saha, Yasin Ibrahim, Cheng Ouyang 외

Learning to Defer (L2D) trains autonomous systems to handle straightforward cases while deferring uncertain ones to human experts. Recent advancements in this field have introduced methods that offer flexibility to unsee…

Diversity

Fatigue-Aware Learning to Defer via Constrained Optimisation

2026-04-01 · Zheng Zhang, Cuong C. Nguyen, David Rosewarne, Kevin Wells 외 arxiv

Learning to defer (L2D) enables human-AI cooperation by deciding when an AI system should act autonomously or defer to a human expert. Existing L2D methods, however, assume static human performance, contradicting well-es…

DeferredSeg:A Multi-Expert Deferral Framework for Medical Image Segmentation

2026-04-14 · Qiuyu Tian, Haoliang Sun, Yunshan Wang, Yinghuan Shi 외 arxiv

Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit overconfidence or underconfidence, leading to unreliable confidence scores f…

Medical Image Segmentation