paper-with-me

Papers

Mathematics of statistical sequential decision-making: concentration, risk-awareness and modelling in stochastic bandits, with applications to bariatric surgery

2024-05-03 · Patrick Saux

This thesis aims to study some of the mathematical challenges that arise in the analysis of statistical sequential decision-making algorithms for postoperative patients follow-up. Stochastic bandits (multiarmed, contextual) model the learning of a sequence of actions (policy) by an agent in an uncertain environment in order to maximise observed rewards. To learn optimal policies, bandit algorithms have to balance the exploitation of current knowledge and the exploration of uncertain actions. Such algorithms have largely been studied and deployed in industrial applications with large datasets, low-risk decisions and clear modelling assumptions, such as clickthrough rate maximisation in online advertising. By contrast, digital health recommendations call for a whole new paradigm of small samples, risk-averse agents and complex, nonparametric modelling. To this end, we developed new safe, anytime-valid concentration bounds, (Bregman, empirical Chernoff), introduced a new framework for risk-aware contextual bandits (with elicitable risk measures) and analysed a novel class of nonparametric bandit algorithms under weak assumptions (Dirichlet sampling). In addition to the theoretical guarantees, these results are supported by in-depth empirical evidence. Finally, as a first step towards personalised postoperative follow-up recommendations, we developed with medical doctors and surgeons an interpretable machine learning model to predict the long-term weight trajectories of patients after bariatric surgery.

📄 PDF Abstract BibTeX arXiv:2405.01994

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingInterpretable Machine LearningMulti-Armed BanditsSequential Decision Makingvalid

Similar Papers 제목 키워드 기반

Selective Reviews of Bandit Problems in AI via a Statistical View

2024-12-03 · Pengjie Zhou, Haoyu Wei, Huiming Zhang

Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A key subset includes stochastic multi-armed…

Decision MakingDecision Making Under UncertaintyMulti-Armed BanditsReinforcement Learning (RL)+1

Vector-valued self-normalized concentration inequalities beyond sub-Gaussianity

2025-11-05 · Diego Martinez-Taboada, Tomas Gonzalez, Aaditya Ramdas arxiv

The study of self-normalized processes plays a crucial role in a wide range of applications, from sequential decision-making to econometrics. While the behavior of self-normalized concentration has been widely investigat…

Adaptive Concentration Inequalities for Sequential Decision Problems

2016-12-01 · NeurIPS 2016 12 · Shengjia Zhao, Enze Zhou, Ashish Sabharwal, Stefano Ermon

A key challenge in sequential decision problems is to determine how many samples are needed for an agent to make reliable decisions with good probabilistic guarantees. We introduce Hoeffding-like concentration inequali…

Two-sample testing

Probability Tools for Sequential Random Projection

2024-02-16 · Yingru Li

We introduce the first probabilistic framework tailored for sequential random projection, an approach rooted in the challenges of sequential decision-making under uncertainty. The analysis is complicated by the sequentia…

Decision MakingDecision Making Under UncertaintyLEMMASequential Decision Making

Linear Stochastic Bandits over a Bit-Constrained Channel

2022-03-02 · Aritra Mitra, Hamed Hassani, George J. Pappas

One of the primary challenges in large-scale distributed learning stems from stringent communication constraints. While several recent works address this challenge for static optimization problems, sequential decision-ma…

Decision MakingDecision Making Under UncertaintySequential Decision Making