Weathering Ongoing Uncertainty: Learning and Planning in a Time-Varying Partially Observable Environment
Optimal decision-making presents a significant challenge for autonomous systems operating in uncertain, stochastic and time-varying environments. Environmental variability over time can significantly impact the system's optimal decision making strategy for mission completion. To model such environments, our work combines the previous notion of Time-Varying Markov Decision Processes (TVMDP) with partial observability and introduces Time-Varying Partially Observable Markov Decision Processes (TV-POMDP). We propose a two-pronged approach to accurately estimate and plan within the TV-POMDP: 1) Memory Prioritized State Estimation (MPSE), which leverages weighted memory to provide more accurate time-varying transition estimates; and 2) an MPSE-integrated planning strategy that optimizes long-term rewards while accounting for temporal constraint. We validate the proposed framework and algorithms using simulations and hardware, with robots exploring a partially observable, time-varying environments. Our results demonstrate superior performance over standard methods, highlighting the framework's effectiveness in stochastic, uncertain, time-varying domains.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingState EstimationSimilar Papers 제목 키워드 기반
Learning Real-World Image De-Weathering with Imperfect Supervision
Real-world image de-weathering aims at removing various undesirable weather-related artifacts. Owing to the impossibility of capturing image pairs concurrently, existing real-world de-weathering datasets often exhibit in…
Pseudo LabelPseudo-Label Guided Real-World Image De-weathering: A Learning Framework with Imperfect Supervision
Real-world image de-weathering aims at removingvarious undesirable weather-related artifacts, e.g., rain, snow,and fog. To this end, acquiring ideal training pairs is crucial.Existing real-world datasets are typically co…
Pseudo LabelInitial validation of a soil-based mass-balance approach for empirical monitoring of enhanced rock weathering rates
Enhanced Rock Weathering (ERW) is a promising scalable and cost-effective Carbon Dioxide Removal (CDR) strategy with significant environmental and agronomic co-benefits. A major barrier to large-scale implementation of E…
A Regret Minimization Approach to Iterative Learning Control
We consider the setting of iterative learning control, or model-based policy learning in the presence of uncertain, time-varying dynamics. In this setting, we propose a new performance metric, planning regret, which repl…
Dynamic Real-time Multimodal Routing with Hierarchical Hybrid Planning
We introduce the problem of Dynamic Real-time Multimodal Routing (DREAMR), which requires planning and executing routes under uncertainty for an autonomous agent. The agent has access to a time-varying transit vehicle ne…
Decision MakingDecision Making Under UncertaintySequential Decision Making