paper-with-me

Papers

A Reduction-based Framework for Sequential Decision Making with Delayed Feedback

2023-02-03 · NeurIPS 2023 11 · Yunchang Yang, Han Zhong, Tianhao Wu, Bin Liu, LiWei Wang, Simon S. Du

We study stochastic delayed feedback in general multi-agent sequential decision making, which includes bandits, single-agent Markov decision processes (MDPs), and Markov games (MGs). We propose a novel reduction-based framework, which turns any multi-batched algorithm for sequential decision making with instantaneous feedback into a sample-efficient algorithm that can handle stochastic delays in sequential decision making. By plugging different multi-batched algorithms into our framework, we provide several examples demonstrating that our framework not only matches or improves existing results for bandits, tabular MDPs, and tabular MGs, but also provides the first line of studies on delays in sequential decision making with function approximation. In summary, we provide a complete set of sharp results for multi-agent sequential decision making with delayed feedback.

📄 PDF Abstract BibTeX arXiv:2302.01477

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingSequential Decision Making

Similar Papers 제목 키워드 기반

Online Sequential Decision-Making with Unknown Delays

2024-02-12 · Ping Wu, Heyan Huang, Zhengyang Liu

In the field of online sequential decision-making, we address the problem with delays utilizing the framework of online convex optimization (OCO), where the feedback of a decision can arrive with an unknown delay. Unlike…

Decision MakingSequential Decision Making

Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models

2026-04-20 · Zehua Zang, Xi Wang, Fuchun Sun, Xiao Xu 외 arxiv

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to …

Test-time AdaptationData Augmentation

Delayed Feedback in Generalised Linear Bandits Revisited

2022-07-21 · Benjamin Howson, Ciara Pike-Burke, Sarah Filippi

The stochastic generalised linear bandit is a well-understood model for sequential decision-making problems, with many algorithms achieving near-optimal regret guarantees under immediate feedback. However, the stringent …

Decision MakingSequential Decision Making

Non-stationary Delayed Combinatorial Semi-Bandit with Causally Related Rewards

2023-07-18 · Saeed Ghoorchian, Setareh Maghsudi

Sequential decision-making under uncertainty is often associated with long feedback delays. Such delays degrade the performance of the learning agent in identifying a subset of arms with the optimal collective reward in …

Decision MakingDecision Making Under UncertaintySequential Decision Making

Smart Commander: A Hierarchical Reinforcement Learning Framework for Fleet-Level PHM Decision Optimization

2026-04-08 · Yong Si, Mingfei Lu, Jing Li, Yang Hu 외 arxiv

Decision-making in military aviation Prognostics and Health Management (PHM) faces significant challenges due to the "curse of dimensionality" in large-scale fleet operations, combined with sparse feedback and stochastic…

Hierarchical Reinforcement Learning