The Devil is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection
Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. In this paper, we dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits to a deep excavation of reciprocal information underlying the entire task. We introduce a Dynamic Feature Reflecting Network, named DFR-Net, which contains two novel standalone modules: (i) the Appearance-Localization Feature Reflecting module (ALFR) that first separates taskspecific features and then self-mutually reflects the reciprocal features; (ii) the Dynamic Intra-Trading module (DIT) that adaptively realigns the training processes of various sub-tasks via a self-learning manner. Extensive experiments on the challenging KITTI dataset demonstrate the effectiveness and generalization of DFR-Net. We rank 1st among all the monocular 3D object detectors in the KITTI test set (till March 16th, 2021). The proposed method is also easy to be plug-and-play in many cutting-edge 3D detection frameworks at negligible cost to boost performance. The code will be made publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionAutonomous DrivingMonocular 3D Object DetectionObjectobject-detectionObject DetectionObject LocalizationSelf-LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Reciprocal Transformations for Unsupervised Video Object Segmentation
Unsupervised video object segmentation (UVOS) aims at segmenting the primary objects in videos without any human intervention. Due to the lack of prior knowledge about the primary objects, identifying them from video…
ObjectOptical Flow EstimationSemantic SegmentationUnsupervised Video Object Segmentation+2Reciprocal Nucleopeptides as the Ancestral Darwinian Self-Replicator
Even the simplest organisms are too complex to have spontaneously arisen fully-formed, yet precursors to first life must have emerged ab initio from their environment. A watershed event was the appearance of the first en…
Stopping Rules for Bag-of-Words Image Search and Its Application in Appearance-Based Localization
We propose a technique to improve the search efficiency of the bag-of-words (BoW) method for image retrieval. We introduce a notion of difficulty for the image matching problems and propose methods that reduce the amount…
Image RetrievalQuantizationRetrievalVisual Localization Under Appearance Change: Filtering Approaches
A major focus of current research on place recognition is visual localization for autonomous driving. In this scenario, as cameras will be operating continuously, it is realistic to expect videos as an input to visual lo…
Visual LocalizationVisual Place RecognitionDetector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and reasoning. In these tasks, an MLLM must lo…
Spatio-Temporal Video Grounding