paper-with-me

홈 › Papers

Where am I? Cross-View Geo-localization with Natural Language Descriptions

2024-12-22 · Junyan Ye, Honglin Lin, Leyan Ou, Dairong Chen, ZiHao Wang, Qi Zhu, Conghui He, Weijia Li

Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ .

📄 PDF Abstract BibTeX arXiv:2412.17007

Code (1)

yejy53/ep-bev pytorch

Tasks

geo-localizationImage RetrievalRetrievalText GenerationText Retrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

2026-08-27 · Jiahao Wen, Hang Yu, Zhedong Zheng arxiv

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended…

A Survey on Video Moment Localization

2023-06-13 · Meng Liu, Liqiang Nie, Yunxiao Wang, Meng Wang 외

Video moment localization, also known as video moment retrieval, aiming to search a target segment within a video described by a given natural language query. Beyond the task of temporal action localization whereby the t…

Action LocalizationMoment RetrievalRetrievalSurvey+1

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

2026-07-14 · Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu 외 arxiv

Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally…

Pose Estimation

Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN

2024-08-13 · Hao Li, Fabian Deuser, Wenping Yina, Xuanshu Luo 외

Nature disasters play a key role in shaping human-urban infrastructure interactions. Effective and efficient response to natural disasters is essential for building resilience and a sustainable urban environment. Two typ…

Contrastive LearningDisaster Response

ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

2026-05-29 · Kaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou 외 arxiv

Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a diagnost…

Common Sense ReasoningSpatial Reasoning