Learning under Invariable Bayesian Safety
A recent body of work addresses safety constraints in explore-and-exploit systems. Such constraints arise where, for example, exploration is carried out by individuals whose welfare should be balanced with overall welfare. In this paper, we adopt a model inspired by recent work on a bandit-like setting for recommendations. We contribute to this line of literature by introducing a safety constraint that should be respected in every round and determines that the expected value in each round is above a given threshold. Due to our modeling, the safe explore-and-exploit policy deserves careful planning, or otherwise, it will lead to sub-optimal welfare. We devise an asymptotically optimal algorithm for the setting and analyze its instance-dependent convergence rate.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Consistent Assistant Domains Transformer for Source-free Domain Adaptation
Source-free domain adaptation (SFDA) aims to address the challenge of adapting to a target domain without accessing the source domain directly. However, due to the inaccessibility of source domain data, deterministic inv…
Source-Free Domain AdaptationQuantitative Metrics for Benchmarking Human-Aware Robot Navigation
Social robots have recently gained popularity, and many human-aware navigation approaches have emerged. This work presents a comprehensive benchmark for quantitatively assessing robot navigation methods. As an automated …
BenchmarkingRobot NavigationLipschitz Safe Bayesian Optimization for Automotive Control
Controller tuning is a labor-intensive process that requires human intervention and expert knowledge. Bayesian optimization has been applied successfully in different fields to automate this process. However, when tuning…
Bayesian OptimizationBayesian Learning-Based Adaptive Control for Safety Critical Systems
Deep learning has enjoyed much recent success, and applying state-of-the-art model learning methods to controls is an exciting prospect. However, there is a strong reluctance to use these methods on safety-critical syste…
Autonomous VehiclesBayesian InferenceGaussian ProcessesInterpretable Machine LearningA Measurement of the Kuiper Belt's Mean Plane From Objects Classified By Machine Learning
Mean plane measurements of the Kuiper Belt from observational data are of interest for their potential to test dynamical models of the solar system. Recent measurements have yielded inconsistent results. Here we report a…