paper-with-me

홈 › Papers

Text2Loc: 3D Point Cloud Localization from Natural Language

2023-11-27 · CVPR 2024 1 · Yan Xia, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to $2\times$ over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at \url{https://yan-xia.github.io/projects/text2loc/}.

📄 PDF Abstract BibTeX arXiv:2311.15977

Code (1)

kevin301342/cmmloc pytorch

Tasks

Contrastive LearningVisual Place Recognition

Methods 이 논문이 사용한 방법론

HINT An unsupervised approach for identifying Hierarchical Information Threads by analysing the network of related articles in a collection. In particular, HINT leverages article…

Similar Papers 제목 키워드 기반

Text2Pos: Text-to-Point-Cloud Cross-Modal Localization

2022-03-28 · CVPR 2022 1 · Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixe

Natural language-based communication with mobile devices and home appliances is becoming increasingly popular and has the potential to become natural for communicating with mobile robots in the future. Towards this goal,…

Visual Place Recognition

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

2026-03-10 · Shuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao 외 arxiv

Text-to-point-cloud (T2P) localization aims to infer precise spatial positions within 3D point cloud maps from natural language descriptions, reflecting how humans perceive and communicate spatial layouts through languag…

Spatial ReasoningPoint Clouds

MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms

2024-08-28 · Tianyi Shang, Zhenyu Li, Pengjie Xu, Jinwei Qiao

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcom…

Cross-modal place recognitionMambaVisual Place Recognition

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

2025-11-19 · Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques 외 arxiv

We tackle the problem of localizing 3D point cloud submaps using complex and diverse natural language descriptions, and present Text2Loc++, a novel neural network designed for effective cross-modal alignment between lang…

Contrastive LearningPoint Clouds

Text to Point Cloud Localization with Relation-Enhanced Transformer

2023-01-13 · Guangzhi Wang, Hehe Fan, Mohan Kankanhalli

Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on the text-to-point-cloud cross…

Natural Language QueriesRelation