Instance-aware multi-object self-supervision for monocular depth prediction
This paper proposes a self-supervised monocular image-to-depth prediction framework that is trained with an end-to-end photometric loss that handles not only 6-DOF camera motion but also 6-DOF moving object instances. Self-supervision is performed by warping the images across a video sequence using depth and scene motion including object instances. One novelty of the proposed method is the use of the multi-head attention of the transformer network that matches moving objects across time and models their interaction and dynamics. This enables accurate and robust pose estimation for each object instance. Most image-to-depth predication frameworks make the assumption of rigid scenes, which largely degrades their performance with respect to dynamic objects. Only a few SOTA papers have accounted for dynamic objects. The proposed method is shown to outperform these methods on standard benchmarks and the impact of the dynamic motion on these benchmarks is exposed. Furthermore, the proposed image-to-depth prediction framework is also shown to be competitive with SOTA video-to-depth prediction frameworks.
Code (0)
등록된 구현이 없습니다.
Tasks
Depth EstimationDepth PredictionObjectPose EstimationPredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Object-Aware Self-supervised Multi-Label Learning
Multi-label Learning on Image data has been widely exploited with deep learning models. However, supervised training on deep CNN models often cannot discover sufficient discriminative features for classification. As a re…
Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-Label LearningObjectHunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation
Recent weakly supervised semantic segmentation (WSSS) methods strive to incorporate contextual knowledge to improve the completeness of class activation maps (CAM). In this work, we argue that the knowledge bias between …
Learning TheorySemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic SegmentationAre Labels Needed for Incremental Instance Learning?
In this paper, we learn to classify visual object instances, incrementally and via self-supervision (self-incremental). Our learner observes a single instance at a time, which is then discarded from the dataset. Incremen…
ObjectPanoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation
We present Panoptic Neural Fields (PNF), an object-aware neural scene representation that decomposes a scene into a set of objects (things) and background (stuff). Each object is represented by an oriented 3D bounding bo…
2D Panoptic Segmentation3D scene EditingDepth EstimationDepth Prediction+3Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection
Weakly supervised learning has emerged as a compelling tool for object detection by reducing the need for strong supervision during training. However, major challenges remain: (1) differentiation of object instances can …
Objectobject-detectionObject DetectionVideo Object Detection+2