paper-with-me

홈 › Papers

Text2Pos: Text-to-Point-Cloud Cross-Modal Localization

2022-03-28 · CVPR 2022 1 · Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixe

Natural language-based communication with mobile devices and home appliances is becoming increasingly popular and has the potential to become natural for communicating with mobile robots in the future. Towards this goal, we investigate cross-modal text-to-point-cloud localization that will allow us to specify, for example, a vehicle pick-up or goods delivery location. In particular, we propose Text2Pos, a cross-modal localization module that learns to align textual descriptions with localization cues in a coarse- to-fine manner. Given a point cloud of the environment, Text2Pos locates a position that is specified via a natural language-based description of the immediate surroundings. To train Text2Pos and study its performance, we construct KITTI360Pose, the first dataset for this task based on the recently introduced KITTI360 dataset. Our experiments show that we can localize 65% of textual queries within 15m distance to query locations for top-10 retrieved locations. This is a starting point that we hope will spark future developments towards language-based navigation.

📄 PDF Abstract BibTeX arXiv:2203.15125

Code (2)

CV4RA/MambaPlace pytorch
kevin301342/cmmloc pytorch

Tasks

Visual Place Recognition

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms

2024-08-28 · Tianyi Shang, Zhenyu Li, Pengjie Xu, Jinwei Qiao

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcom…

Cross-modal place recognitionMambaVisual Place Recognition

PVContext: Hybrid Context Model for Point Cloud Compression

2024-09-19 · Guoqing Zhang, Wenbo Zhao, Jian Liu, Yuanchao Bai 외

Efficient storage of large-scale point cloud data has become increasingly challenging due to advancements in scanning technology. Recent deep learning techniques have revolutionized this field; However, most existing app…

model

Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding

2025-01-04 · Yingjie Liu, Pengyu Zhang, Ziyao He, Mingsong Chen 외

Hyperbolic spaces allow for more efficient modeling of complex, hierarchical structures, which is particularly beneficial in tasks involving multi-modal data. Although hyperbolic geometries have been proven effective for…

Contrastive Learning

Text-Driven Cross-Modal Place Recognition Method for Remote Sensing Localization

2025-03-23 · Tianyi Shang, Zhenyu Li, Pengjie Xu, ZhaoJun Deng 외

Environment description-based localization in large-scale point cloud maps constructed through remote sensing is critically significant for the advancement of large-scale autonomous systems, such as delivery robots opera…

Cross-modal place recognition

Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction

2025-04-02 · Junlong Ren, Hao Wang

Cross-modal 3D retrieval is a critical yet challenging task, aiming to achieve bi-directional retrieval between 3D and text modalities. Current methods predominantly rely on a certain 3D representation (e.g., point cloud…

Retrieval