Self-Supervised Simultaneous Multi-Step Prediction of Road Dynamics and Cost Map
While supervised learning is widely used for perception modules in conventional autonomous driving solutions, scalability is hindered by the huge amount of data labeling needed. In contrast, while end-to-end architectures do not require labeled data and are potentially more scalable, interpretability is sacrificed. We introduce a novel architecture that is trained in a fully self-supervised fashion for simultaneous multi-step prediction of space-time cost map and road dynamics. Our solution replaces the manually designed cost function for motion planning with a learned high dimensional cost map that is naturally interpretable and allows diverse contextual information to be integrated without manual data labeling. Experiments on real world driving data show that our solution leads to lower number of collisions and road violations in long planning horizons in comparison to baselines, demonstrating the feasibility of fully self-supervised prediction without sacrificing either scalability or interpretability.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingMotion PlanningSimilar Papers 제목 키워드 기반
Self-supervised Pretraining and Finetuning for Monocular Depth and Visual Odometry
For the task of simultaneous monocular depth and visual odometry estimation, we propose learning self-supervised transformer-based models in two steps. Our first step consists in a generic pretraining to learn 3D geometr…
3D geometryDepth EstimationDepth PredictionVisual OdometryProphetNet: Predicting Future N-gram for Sequence-to-SequencePre-training
This paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism. I…
Abstractive Text SummarizationPredictionQuestion GenerationQuestion-GenerationProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
This paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism. I…
Abstractive Text SummarizationPredictionQuestion GenerationQuestion-Generation+1Self-supervised Learning for Single View Depth and Surface Normal Estimation
In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single image. In contrast to most existing framewo…
Depth EstimationDepth PredictionMonocular Depth EstimationSelf-Supervised Learning+1Self-supervised Contrastive Learning of Multi-view Facial Expressions
Facial expression recognition (FER) has emerged as an important component of human-computer interaction systems. Despite recent advancements in FER, performance often drops significantly for non-frontal facial images. We…
Contrastive LearningFacial Expression RecognitionFacial Expression Recognition (FER)