paper-with-me

Papers

FindView: Precise Target View Localization Task for Look Around Agents

2023-03-16 · Haruya Ishikawa, Yoshimitsu Aoki

With the increase in demands for service robots and automated inspection, agents need to localize in its surrounding environment to achieve more natural communication with humans by shared contexts. In this work, we propose a novel but straightforward task of precise target view localization for look around agents called the FindView task. This task imitates the movements of PTZ cameras or user interfaces for 360 degree mediums, where the observer must "look around" to find a view that exactly matches the target. To solve this task, we introduce a rule-based agent that heuristically finds the optimal view and a policy learning agent that employs reinforcement learning to learn by interacting with the 360 degree scene. Through extensive evaluations and benchmarks, we conclude that learned methods have many advantages, in particular precise localization that is robust to corruption and can be easily deployed in novel scenes.

📄 PDF Abstract BibTeX arXiv:2303.09054

Code (1)

haruishi43/look_around 공식 구현

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information

2024-07-22 · Luca Di Giammarino, Boyang Sun, Giorgio Grisetti, Marc Pollefeys 외

Accurate localization in diverse environments is a fundamental challenge in computer vision and robotics. The task involves determining a sensor's precise position and orientation, typically a camera, within a given spac…

Adaloss: Adaptive Loss Function for Landmark Localization

2019-08-02 · Brian Teixeira, Birgi Tamersoy, Vivek Singh, Ankur Kapoor

Landmark localization is a challenging problem in computer vision with a multitude of applications. Recent deep learning based methods have shown improved results by regressing likelihood maps instead of regressing the c…

Facial Landmark Detection

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

2025-03-28 · Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan 외

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-gra…

3D visual groundingAutonomous DrivingScene UnderstandingVisual Grounding

University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization

2020-02-27 · Zhedong Zheng, Yunchao Wei, Yi Yang

We consider the problem of cross-view geo-localization. The primary challenge of this task is to learn the robust feature against large viewpoint changes. Existing benchmarks can help, but are limited in the number of vi…

Drone navigationDrone-view target localizationgeo-localizationImage-Based Localization+1

CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation

2026-06-03 · Yurim Jeon, Dongseong Seo, Seung-Woo Seo arxiv

Cross-view geo-localization estimates the geographic location of a ground image by matching it against an aerial image database. Existing methods tackle this through either large-scale retrieval or precise pose estimatio…

Pose Estimation