paper-with-me

홈 › Papers

LaVPR: Benchmarking Language and Vision for Place Recognition

2026-02-03 · Ofer Idan, Dan Badur, Yosi Keller, Yoli Shavit arxiv

Visual Place Recognition (VPR) often fails under extreme environmental changes and perceptual aliasing. Beyond these limitations, standard systems cannot perform 'blind' localization from verbal descriptions alone, a capability critical for applications such as emergency response. To address these challenges, we introduce LaVPR, a large-scale benchmark that extends existing VPR datasets with over 650,000 rich natural-language descriptions. Using LaVPR, we investigate two paradigms: Multi-Modal Fusion for enhanced robustness and Cross-Modal Retrieval for language-based localization. Our results show that language descriptions yield consistent gains in visually degraded conditions, with the most significant impact on smaller backbones. Notably, adding language allows compact models to rival the performance of much larger vision-only architectures. For cross-modal retrieval, we establish a baseline using Low-Rank Adaptation (LoRA) and Multi-Similarity loss, which substantially outperforms standard contrastive methods across vision-language models. Ultimately, LaVPR enables a new class of localization systems that are both resilient to real-world stochasticity and practical for resource-constrained deployment. Our dataset and code are available at https://github.com/oferidan1/LaVPR

📄 PDF Abstract BibTeX arXiv:2602.03253

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Place RecognitionCross-Modal Retrieval

Similar Papers 제목 키워드 기반

SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition

2025-02-23 · Feng Lu, Tong Jin, Xiangyuan Lan, Lijun Zhang 외

Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a novel method to realize seamless adaptat…

Deep HashingGPURe-RankingRetrieval+1

Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition

2024-02-22 · Feng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong 외

Recent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts h…

Re-RankingVisual Place Recognition

ALTO: A Large-Scale Dataset for UAV Visual Place Recognition and Localization

2022-07-19 · Ivan Cisneros, Peng Yin, Ji Zhang, Howie Choset 외

We present the ALTO dataset, a vision-focused dataset for the development and benchmarking of Visual Place Recognition and Localization methods for Unmanned Aerial Vehicles. The dataset is composed of two long (approxima…

BenchmarkingImage RegistrationVisual OdometryVisual Place Recognition

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

2026-08-02 · Timothee Mickus, Claudio Savelli, Eduardo Calò, Emilio Raimond 외 arxiv

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make b…

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

2025-02-10 · Sankalp Nagaonkar, Augustya Sharma, Ashish Choithani, Ashutosh Trivedi

This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a curated dataset containing 1,477 manual…

BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)