paper-with-me

홈 › Papers

Online Learning and Decision-Making under Generalized Linear Model with High-Dimensional Data

2018-12-07 · Xue Wang, Mike Mingcheng Wei, Tao Yao

We propose a minimax concave penalized multi-armed bandit algorithm under generalized linear model (G-MCP-Bandit) for a decision-maker facing high-dimensional data in an online learning and decision-making process. We demonstrate that the G-MCP-Bandit algorithm asymptotically achieves the optimal cumulative regret in the sample size dimension T , O(log T), and further attains a tight bound in the covariate dimension d, O(log d). In addition, we develop a linear approximation method, the 2-step weighted Lasso procedure, to identify the MCP estimator for the G-MCP-Bandit algorithm under non-iid samples. Under this procedure, the MCP estimator matches the oracle estimator with high probability and converges to the true parameters with the optimal convergence rate. Finally, through experiments based on synthetic data and two real datasets (warfarin dosing dataset and Tencent search advertising dataset), we show that the G-MCP-Bandit algorithm outperforms other benchmark algorithms, especially when there is a high level of data sparsity or the decision set is large.

📄 PDF Abstract BibTeX arXiv:1812.02962

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

GEC: A Unified Framework for Interactive Decision Making in MDP, POMDP, and Beyond

2022-11-03 · Han Zhong, Wei Xiong, Sirui Zheng, LiWei Wang 외

We study sample efficient reinforcement learning (RL) under the general framework of interactive decision making, which includes Markov decision process (MDP), partially observable Markov decision process (POMDP), and pr…

Decision MakingReinforcement Learning (RL)

Generalized Linear Bandits with Local Differential Privacy

2021-06-07 · NeurIPS 2021 12 · Yuxuan Han, Zhipeng Liang, Yang Wang, Jiheng Zhang

Contextual bandit algorithms are useful in personalized online decision-making. However, many applications such as personalized medicine and online advertising require the utilization of individual-specific information f…

Decision MakingMulti-Armed Bandits

A General Framework for Sequential Decision-Making under Adaptivity Constraints

2023-06-26 · Nuoya Xiong, Zhaoran Wang, Zhuoran Yang

We take the first step in studying general sequential decision-making under two adaptivity constraints: rare policy switch and batch learning. First, we provide a general class called the Eluder Condition class, which in…

Decision MakingSequential Decision Making

Decision Making with Linear Constraints on Probabilities

2013-03-27 · Michael Pittarelli

Techniques for decision making with knowledge of linear constraints on condition probabilities are examined. These constraints arise naturally in many situations: upper and lower condition probabilities are known; an ord…

Decision Making

Likelihood Ratio Confidence Sets for Sequential Decision Making

2023-11-08 · NeurIPS 2023 11

Certifiable, adaptive uncertainty estimates for unknown quantities are an essential ingredient of sequential decision-making algorithms. Standard approaches rely on problem-dependent concentration results and are limited…

Decision MakingSequential Decision MakingSurvival Analysisvalid