Robust Driving Policy Learning with Guided Meta Reinforcement Learning
Although deep reinforcement learning (DRL) has shown promising results for autonomous navigation in interactive traffic scenarios, existing work typically adopts a fixed behavior policy to control social vehicles in the training environment. This may cause the learned driving policy to overfit the environment, making it difficult to interact well with vehicles with different, unseen behaviors. In this work, we introduce an efficient method to train diverse driving policies for social vehicles as a single meta-policy. By randomizing the interaction-based reward functions of social vehicles, we can generate diverse objectives and efficiently train the meta-policy through guiding policies that achieve specific objectives. We further propose a training strategy to enhance the robustness of the ego vehicle's driving policy using the environment where social vehicles are controlled by the learned meta-policy. Our method successfully learns an ego driving policy that generalizes well to unseen situations with out-of-distribution (OOD) social agents' behaviors in a challenging uncontrolled T-intersection scenario.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous NavigationDeep Reinforcement LearningMeta Reinforcement Learningreinforcement-learningReinforcement LearningSimilar Papers 제목 키워드 기반
Importance Sampling-Guided Meta-Training for Intelligent Agents in Highly Interactive Environments
Training intelligent agents to navigate highly interactive environments presents significant challenges. While guided meta reinforcement learning (RL) approach that first trains a guiding policy to train the ego agent ha…
Meta Reinforcement LearningNavigateReinforcement Learning (RL)Composing Meta-Policies for Autonomous Driving Using Hierarchical Deep Reinforcement Learning
Rather than learning new control policies for each new task, it is possible, when tasks share some structure, to compose a "meta-policy" from previously learned policies. This paper reports results from experiments using…
Autonomous DrivingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1Meta-Reinforcement Learning for Adaptive Autonomous Driving
Reinforcement learning (RL) methods achieved major advances in multiple tasks surpassing human performance. However, most of RL strategies show a certain degree of weakness and may become computationally intractable when…
Autonomous DrivingMeta Reinforcement Learningreinforcement-learningReinforcement Learning+1Quick Learner Automated Vehicle Adapting its Roadmanship to Varying Traffic Cultures with Meta Reinforcement Learning
It is essential for an automated vehicle in the field to perform discretionary lane changes with appropriate roadmanship - driving safely and efficiently without annoying or endangering other road users - under a wide ra…
Deep Reinforcement LearningMeta Reinforcement Learningreinforcement-learningReinforcement Learning+1Guided Online Distillation: Promoting Safe Reinforcement Learning by Offline Demonstration
Safe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints. When learning from scratch, safe RL agents tend to be overly conservative, which impedes exploration an…
Autonomous DrivingDecision Makingreinforcement-learningReinforcement Learning+2