paper-with-me

홈 › Papers

GOMAA-Geo: GOal Modality Agnostic Active Geo-localization

2024-06-04 · Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Chongjie Zhang, Nathan Jacobs, Yevgeniy Vorobeychik

We consider the task of active geo-localization (AGL) in which an agent uses a sequence of visual cues observed during aerial navigation to find a target specified through multiple possible modalities. This could emulate a UAV involved in a search-and-rescue operation navigating through an area, observing a stream of aerial images as it goes. The AGL task is associated with two important challenges. Firstly, an agent must deal with a goal specification in one of multiple modalities (e.g., through a natural language description) while the search cues are provided in other modalities (aerial imagery). The second challenge is limited localization time (e.g., limited battery life, urgency) so that the goal must be localized as efficiently as possible, i.e. the agent must effectively leverage its sequentially observed aerial views when searching for the goal. To address these challenges, we propose GOMAA-Geo - a goal modality agnostic active geo-localization agent - for zero-shot generalization between different goal modalities. Our approach combines cross-modality contrastive learning to align representations across modalities with supervised foundation model pretraining and reinforcement learning to obtain highly effective navigation and localization policies. Through extensive evaluations, we show that GOMAA-Geo outperforms alternative learnable approaches and that it generalizes across datasets - e.g., to disaster-hit areas without seeing a single disaster scenario during training - and goal modalities - e.g., to ground-level imagery or textual descriptions, despite only being trained with goals specified as aerial views. Code and models are publicly available at https://github.com/mvrl/GOMAA-Geo/tree/main.

📄 PDF Abstract BibTeX arXiv:2406.01917

Code (2)

mvrl/gomaa-geo 공식 구현 pytorch
aleksispi/airloc pytorch

Tasks

Contrastive Learninggeo-localizationZero-shot Generalization

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Unifying Vision-Language Representation Space with Single-tower Transformer

2022-11-21 · Jiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim 외

Contrastive learning is a form of distance learning that aims to learn invariant features from two related representations. In this paper, we explore the bold hypothesis that an image and its caption can be simply regard…

Contrastive LearningObject LocalizationRepresentation LearningRetrieval+1

Active Semantic Localization with Graph Neural Embedding

2023-05-10 · Mitsuki Yoshida, Kanji Tanaka, Ryogo Yamamoto, Daiki Iwata

Semantic localization, i.e., robot self-localization with semantic image modality, is critical in recently emerging embodied AI applications (e.g., point-goal navigation, object-goal navigation, vision language navigatio…

CPUDomain AdaptationGraph Neural NetworkSelf-Supervised Learning+2

Modality-Agnostic Topology Aware Localization

2021-12-01 · NeurIPS 2021 12 · Farhad Ghazvinian Zanjani, Ilia Karmanov, Hanno Ackermann, Daniel Dijkman 외

This work presents a data-driven approach for the indoor localization of an observer on a 2D topological map of the environment. State-of-the-art techniques may yield accurate estimates only when they are tailor-made for…

Indoor Localization

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration

2025-07-31 · Li Mi, Manon Bechaz, Zeming Chen, Antoine Bosselut 외 arxiv

Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a go…

Reinforcement Learning

Towards Accurate Active Camera Localization

2020-12-08 · Qihang Fang, Yingda Yin, Qingnan Fan, Fei Xia 외

In this work, we tackle the problem of active camera localization, which controls the camera movements actively to achieve an accurate camera pose. The past solutions are mostly based on Markov Localization, which reduce…

Camera LocalizationCamera Pose EstimationPose EstimationVisual Localization