paper-with-me

홈 › Papers

Online MDP with Transition Prototypes: A Robust Adaptive Approach

2024-12-18 · Shuo Sun, Meng Qi, Zuo-Jun Max Shen

In this work, we consider an online robust Markov Decision Process (MDP) where we have the information of finitely many prototypes of the underlying transition kernel. We consider an adaptively updated ambiguity set of the prototypes and propose an algorithm that efficiently identifies the true underlying transition kernel while guaranteeing the performance of the corresponding robust policy. To be more specific, we provide a sublinear regret of the subsequent optimal robust policy. We also provide an early stopping mechanism and a worst-case performance bound of the value function. In numerical experiments, we demonstrate that our method outperforms existing approaches, particularly in the early stage with limited data. This work contributes to robust MDPs by considering possible prior information about the underlying transition probability and online learning, offering both theoretical insights and practical algorithms for improved decision-making under uncertainty.

📄 PDF Abstract BibTeX arXiv:2412.14075

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDecision Making Under Uncertainty

Methods 이 논문이 사용한 방법론

Early Stopping Early Stopping is a regularization technique for deep neural networks that stops training when parameter updates no longer begin to yield improves on a validation set. In…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Analyzing the coupling process of distributed mixed real-virtual prototypes

2024-01-22 · Peter Baumann, Lars Mikelsons, Oliver Kotte, Dieter Schramm

The ongoing connection and automation of vehicles leads to a closer interaction of the individual vehicle components, which demands for consideration throughout the entire development process. In the design phase, this i…

Asymmetric Adaptation-based Real-time Fault Diagnosis Under Transitional Operating Conditions

2026-05-23 · Hongshuo Zhao, Zeyi Liu, Xiao He arxiv

Data streams in real-world industrial scenarios often contain transitional operating conditions that are uncovered during offline training, leading to significant distribution shifts. To bridge the gap between static off…

Domain GeneralizationTest-time AdaptationFault Diagnosis

Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation

2021-01-26 · CVPR 2021 1 · Pan Zhang, Bo Zhang, Ting Zhang, Dong Chen 외

Self-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are d…

Domain AdaptationImage-to-Image TranslationPseudo LabelSemantic Segmentation+2

Exploring emotional prototypes in a high dimensional TTS latent space

2021-05-05 · Pol van Rijn, Silvan Mertes, Dominik Schiller, Peter M. C. Harrison 외

Recent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psy…

Vocal Bursts Intensity Prediction

Online Tuning for Offline Decentralized Multi-Agent Reinforcement Learning

2021-09-29 · Jiechuan Jiang, Zongqing Lu

Offline reinforcement learning could learn effective policies from a fixed dataset, which is promising in real-world applications. However, in offline decentralized multi-agent reinforcement learning, due to the discrepa…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)