paper-with-me

홈 › Papers

Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs

2026-03-30 · Pei An, Junfeng Ding, Jiaqi Yang, Yulong Wang, Jie Ma, Liangliang Nan arxiv

Image-to-point-cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D-3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are both discriminative and generalizable, leading to severe performance drops in unseen scenarios. We address this challenge by introducing a heterogeneous graph that enables refining both cross-modal features and correspondences within a unified architecture. The proposed graph represents a mapping between segmented 2D and 3D regions, which enhances cross-modal feature interaction and thus improves feature discriminability. In addition, modeling the consistency among vertices and edges within the graph enables pruning of unreliable correspondences. Building on these insights, we propose a heterogeneous graph embedded I2P registration method, termed Hg-I2P. It learns a heterogeneous graph by mining multi-path feature relationships, adapts features under the guidance of heterogeneous edges, and prunes correspondences using graph-based projection consistency. Experiments on six indoor and outdoor benchmarks under cross-domain setups demonstrate that Hg-I2P significantly outperforms existing methods in both generalization and accuracy. Code is released on https://github.com/anpei96/hg-i2p-demo.

📄 PDF Abstract BibTeX arXiv:2603.27969

Code (0)

등록된 구현이 없습니다.

Tasks

Point Clouds

Similar Papers 제목 키워드 기반

Domain Adaptation for Vehicle Detection from Bird's Eye View LiDAR Point Cloud Data

2019-05-22 · Khaled Saleh, Ahmed Abobakr, Mohammed Attia, Julie Iskander 외

Point cloud data from 3D LiDAR sensors are one of the most crucial sensor modalities for versatile safety-critical applications such as self-driving vehicles. Since the annotations of point cloud data is an expensive and…

Domain AdaptationUnsupervised Domain Adaptationvehicle detection

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

2026-01-27 · Jiarun Liu, Qifeng Chen, Yiru Zhao, Minghua Liu 외 arxiv

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g.,…

Cross-Modal RetrievalScene RecognitionPoint Clouds

Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning

2024-02-04 · Haoyi Zhu, Yating Wang, Di Huang, Weicai Ye 외

In robot learning, the observation space is crucial due to the distinct characteristics of different modalities, which can potentially become a bottleneck alongside policy design. In this study, we explore the influence …

Contact-rich ManipulationZero-shot Generalization

Point Cloud-Assisted Neural Image Compression

2024-12-16 · Ziqun Li, Qi Zhang, Xiaofeng Huang, Zhao Wang 외

High-efficient image compression is a critical requirement. In several scenarios where multiple modalities of data are captured by different sensors, the auxiliary information from other modalities are not fully leverage…

Autonomous DrivingImage Compression

LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework

2025-05-30 · Xin Kang, Zihan Zheng, Lei Chu, Yue Gao 외

We present LTM3D, a Latent Token space Modeling framework for conditional 3D shape generation that integrates the strengths of diffusion and auto-regressive (AR) models. While diffusion-based methods effectively model co…

3D Generation3D Shape Generation