paper-with-me

Papers

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

2025-09-30 · Shunpeng Chen, Changwei Wang, Rongtao Xu, Xingtian Pei, Yukun Song, Jinzhou Lin, Wenhao Xu, Jingyi Zhang, Li Guo, Shibiao Xu arxiv

Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior methods focus on descriptor fine-tuning or fixed sampling strategies yet neglect the dynamic interplay between spatial context and visual similarity during training. We present SAGE (Spatial-visual Adaptive Graph Exploration), a unified training pipeline that enhances granular spatial-visual discrimination by jointly improving local feature aggregation, organize samples during training, and hard sample mining. We introduce a lightweight Soft Probing module that learns residual weights from training data for patch descriptors before bilinear aggregation, boosting distinctive local cues. During training we reconstruct an online geo-visual graph that fuses geographic proximity and current visual similarity so that candidate neighborhoods reflect the evolving embedding landscape. To concentrate learning on the most informative place neighborhoods, we seed clusters from high-affinity anchors and iteratively expand them with a greedy weighted clique expansion sampler. Implemented with a frozen DINOv2 backbone and parameter-efficient fine-tuning, SAGE achieves SOTA across eight benchmarks. Notably, our method obtains 100% Recall@10 on SPED only using 4096D global descriptors. The code and model are available at https://github.com/chenshunpeng/SAGE.

📄 PDF Abstract BibTeX arXiv:2509.25723

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningVisual Place Recognition

Similar Papers 제목 키워드 기반

GestureLens: Visual Analysis of Gestures in Presentation Videos

2022-04-19 · Haipeng Zeng, Xingbo Wang, Yong Wang, Aoyu Wu 외

Appropriate gestures can enhance message delivery and audience engagement in both daily communication and public presentations. In this paper, we contribute a visual analytic approach that assists professional public spe…

InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization

2025-08-07 · Yuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li 외 arxiv

The emergence of Multimodal Large Language Models (MLLMs) has propelled the development of autonomous agents that operate on Graphical User Interfaces (GUIs) using pure visual input. A fundamental challenge is robustly g…

Reinforcement LearningAnswer Generation

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

2026-05-25 · Dong-Hee Kim, Reuben Tan, Donghyun Kim arxiv

Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search tasks, their role in more complex visual r…

Visual Question AnsweringSpatial ReasoningVisual Reasoning

A Natural-language-based Visual Query Approach of Uncertain Human Trajectories

2019-08-01 · Zhaosong Huang, Ye Zhao, Wei Chen, Shengjie Gao 외

Visual querying is essential for interactively exploring massive trajectory data. However, the data uncertainty imposes profound challenges to fulfill advanced analytics requirements. On the one hand, many underlying dat…

Sentence

SPACE: 3D Spatial Co-operation and Exploration Framework for Robust Mapping and Coverage with Multi-Robot Systems

2024-11-04 · Sai Krishna Ghanta, Ramviyas Parasuraman

In indoor environments, multi-robot visual (RGB-D) mapping and exploration hold immense potential for application in domains such as domestic service and logistics, where deploying multiple robots in the same environment…

Point cloud reconstruction