paper-with-me

홈 › Papers

$L^3$:Scene-agnostic Visual Localization in the Wild

2026-03-09 · Yu Zhang, Muhua Zhu, Yifei Xue, Tie Ji, Yizhen Lao arxiv

Standard visual localization methods typically require offline pre-processing of scenes to obtain 3D structural information for better performance. This inevitably introduces additional computational and time costs, as well as the overhead of storing scene representations. Can we visually localize in a wild scene without any off-line preprocessing step? In this paper, we leverage the online inference capabilities of feed-forward 3D reconstruction networks to propose a novel map-free visual localization framework $L^3$. Specifically, by performing direct online 3D reconstruction on RGB images, followed by two-stage metric scale recovery and pose refinement based on 2D-3D correspondences, $L^3$ achieves high accuracy without the need to pre-build or store any offline scene representations. Extensive experiments demonstrate $L^3$ not only that the performance is comparable to state-of-the-art solutions on various benchmarks, but also that it exhibits significantly superior robustness in sparse scenes (fewer reference images per scene).

📄 PDF Abstract BibTeX arXiv:2603.07937

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Localization3D Reconstruction

Similar Papers 제목 키워드 기반

Visual Sound Localization in the Wild by Cross-Modal Interference Erasing

2022-02-13 · Xian Liu, Rui Qian, Hang Zhou, Di Hu 외

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sou…

Sound Source Localization

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

2023-04-12 · Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng 외

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a…

3D visual groundingAutonomous DrivingObject LocalizationVisual Grounding

TextPlace: Visual Place Recognition and Topological Localization Through Reading Scene Texts

2019-10-01 · ICCV 2019 10 · Ziyang Hong, Yvan Petillot, David Lane, Yishu Miao 외

Visual place recognition is a fundamental problem for many vision based applications. Sparse feature and deep learning based methods have been successful and dominant over the decade. However, most of them do not explici…

Visual Place Recognition

JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

2026-03-30 · Koki Maeda, Naoaki Okazaki arxiv

Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese…

Key Information ExtractionVisual Question Answering

SceneAligner: 3D-Grounded Floorplan Localization in the Wild

2026-05-21 · Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor arxiv

Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observati…