RoRD: Rotation-Robust Descriptors and Orthographic Views for Local Feature Matching
The use of local detectors and descriptors in typical computer vision pipelines work well until variations in viewpoint and appearance change become extreme. Past research in this area has typically focused on one of two approaches to this challenge: the use of projections into spaces more suitable for feature matching under extreme viewpoint changes, and attempting to learn features that are inherently more robust to viewpoint change. In this paper, we present a novel framework that combines learning of invariant descriptors through data augmentation and orthographic viewpoint projection. We propose rotation-robust local descriptors, learnt through training data augmentation based on rotation homographies, and a correspondence ensemble technique that combines vanilla feature correspondences with those obtained through rotation-robust features. Using a range of benchmark datasets as well as contributing a new bespoke dataset for this research domain, we evaluate the effectiveness of the proposed approach on key tasks including pose estimation and visual place recognition. Our system outperforms a range of baseline and state-of-the-art techniques, including enabling higher levels of place recognition precision across opposing place viewpoints and achieves practically-useful performance levels even under extreme viewpoint changes.
Code (1)
Tasks
Pose EstimationVisual Place RecognitionSimilar Papers 제목 키워드 기반
RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration
Successful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-variant backbone whose performance declines un…
Point Cloud Registration3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial int…
Spatial ReasoningPPF-FoldNet: Unsupervised Learning of Rotation Invariant 3D Local Descriptors
We present PPF-FoldNet for unsupervised learning of 3D local descriptors on pure point cloud geometry. Based on the folding-based auto-encoding of well known point pair features, PPF-FoldNet offers many desirable propert…
Point Cloud RegistrationLearning general and distinctive 3D local deep descriptors for point cloud registration
An effective 3D descriptor should be invariant to different geometric transformations, such as scale and rotation, robust to occlusions and clutter, and capable of generalising to different application domains. We presen…
Image to Point Cloud RegistrationPoint Cloud RegistrationVNI-Net: Vector Neurons-based Rotation-Invariant Descriptor for LiDAR Place Recognition
LiDAR-based place recognition plays a crucial role in Simultaneous Localization and Mapping (SLAM) and LiDAR localization. Despite the emergence of various deep learning-based and hand-crafting-based methods, rotation-in…
Computational EfficiencySimultaneous Localization and Mapping