Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration
Bridging 2D and 3D sensor modalities is critical for robust perception in autonomous systems. However, image-to-point cloud (I2P) registration remains challenging due to the semantic-geometric gap between texture-rich but depth-ambiguous images and sparse yet metrically precise point clouds, as well as the tendency of existing methods to converge to local optima. To overcome these limitations, we introduce CrossI2P, a self-supervised framework that unifies cross-modal learning and two-stage registration in a single end-to-end pipeline. First, we learn a geometric-semantic fused embedding space via dual-path contrastive learning, enabling annotation-free, bidirectional alignment of 2D textures and 3D structures. Second, we adopt a coarse-to-fine registration paradigm: a global stage establishes superpoint-superpixel correspondences through joint intra-modal context and cross-modal interaction modeling, followed by a geometry-constrained point-level refinement for precise registration. Third, we employ a dynamic training mechanism with gradient normalization to balance losses for feature alignment, correspondence refinement, and pose estimation. Extensive experiments demonstrate that CrossI2P outperforms state-of-the-art methods by 23.7% on the KITTI Odometry benchmark and by 37.9% on nuScenes, significantly improving both accuracy and robustness.
Code (0)
등록된 구현이 없습니다.
Tasks
Point Cloud RegistrationContrastive LearningPose EstimationPoint CloudsSimilar Papers 제목 키워드 기반
PointCMC: Cross-Modal Multi-Scale Correspondences Learning for Point Cloud Understanding
Some self-supervised cross-modal learning approaches have recently demonstrated the potential of image signals for enhancing point cloud representation. However, it remains a question on how to directly model cross-modal…
3D Object ClassificationRepresentation LearningCrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
This paper introduces a novel approach named CrossVideo, which aims to enhance self-supervised cross-modal contrastive learning in the field of point cloud video understanding. Traditional supervised learning methods enc…
Contrastive Learningpoint cloud video understandingSelf-Supervised LearningVideo UnderstandingCrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud Understanding
Manual annotation of large-scale point cloud dataset for varying tasks such as 3D object classification, segmentation and detection is often laborious owing to the irregular structure of point clouds. Self-supervised lea…
3D Object Classification3D Point Cloud Linear ClassificationContrastive LearningFew-Shot 3D Point Cloud Classification+1Self-Supervised Modality-Invariant and Modality-Specific Feature Learning for 3D Objects
While most existing self-supervised 3D feature learning methods mainly focus on point cloud data, this paper explores the inherent multimodal attributes of 3D objects. We propose to jointly learn effective features from …
3D Object RecognitionCross-Modal RetrievalObject RecognitionRetrievalSelf-supervised Feature Learning by Cross-modality and Cross-view Correspondences
The success of supervised learning requires large-scale ground truth labels which are very expensive, time-consuming, or may need special skills to annotate. To address this issue, many self- or un-supervised methods are…
3D Part Segmentation3D Shape Classification3D Shape Recognition3D Shape Retrieval+2