paper-with-me

Papers

Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization

2026-06-23 · Zhaopeng Cui, Jiarui Hu, Jingbo Liu, Boming Zhao, Xiyue Guo, Boyin Feng, Haocheng Peng, Yujun Shen, Hujun Bao, Guofeng Zhang arxiv

Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted to low-level vision schemes, struggling to perceive scene semantics and compositions, which limits both interpretability and applicability. In this paper, we explore the issue of how to organize rich object information in a scene, including semantics, layout, and geometry, into a structured map representation, thereby utilizing object units exclusively to drive the camera relocalization task. To this end, we propose OpenReLoc, a camera relocalization system designed to provide scene understanding and accurate pose estimation capabilities. Leveraging recent foundation models, we first introduce a multi-modal mechanism to integrate open-vocabulary semantic knowledge for effective 2D-3D object matching. Additionally, we design object-oriented reference frames as position priors, paired with a reference frame selection strategy based on the Distance-IoU (DIOU), enabling extension to scalable scenes. Moreover, to ensure stable and accurate pose optimization, we also propose a dual-path 2D Iterative Closest Pixel loss guided by object shape. Experimental results demonstrate that OpenReLoc achieves superior relocalization recall and accuracy across various datasets. Our source code will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2606.24767

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingPose Estimation

Similar Papers 제목 키워드 기반

OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation

2026-08-14 · Dinh Tuan Nguyen, Anh Dao, Phuong Nam Dang, Quan-Dung Pham 외 arxiv

Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove …

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

2024-09-27 · Ayca Takmaz, Alexandros Delitzas, Robert W. Sumner, Francis Engelmann 외

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances but …

3D Instance Segmentation3D Part SegmentationInstance SegmentationObject+3

T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability

2026-04-20 · Savya Khosla, Sethuraman T, Aryan Chadha, Alex Schwing 외 arxiv

Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-vocabulary semantic segmentation; and (2) h…

Semantic SegmentationObject LocalizationImage RetrievalScene Parsing

Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation

2024-01-18 · Zesen Cheng, Kehan Li, Hao Li, Peng Jin 외

Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage image-te…

Instance SegmentationSemantic SegmentationVideo Instance Segmentation

Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection

2022-07-07 · Hanoona Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman Khan 외

Existing open-vocabulary object detectors typically enlarge their vocabulary sizes by leveraging different forms of weak supervision. This helps generalize to novel objects at inference. Two popular forms of weak-supervi…

ObjectOpen Vocabulary Attribute DetectionOpen Vocabulary Object DetectionZero-Shot Object Detection