paper-with-me

홈 › Papers

AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

2025-08-14 · Shixiong Xu, Chenghao Zhang, Lubin Fan, Yuan Zhou, Bin Fan, Shiming Xiang, Gaofeng Meng, Jieping Ye arxiv

Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively.

📄 PDF Abstract BibTeX arXiv:2508.10667

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

2026-04-02 · Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing 외 arxiv

Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we …

Visual Question AnsweringRepresentation LearningScene ClassificationScene Understanding

Discriminative Probing and Tuning for Text-to-Image Generation

2024-03-07 · CVPR 2024 1 · Leigang Qu, Wenjie Wang, Yongqi Li, Hanwang Zhang 외

Despite advancements in text-to-image generation (T2I), prior methods often face text-image misalignment problems such as relation confusion in generated images. Existing solutions involve cross-attention manipulation fo…

Image GenerationText to Image GenerationText-to-Image Generation

Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification

2025-09-22 · Sheng Huang, Jiexuan Yan, Beiyan Liu, Bo Liu 외 arxiv

Real-world datasets often exhibit class imbalance across multiple categories, manifesting as long-tailed distributions and few-shot scenarios. This is especially challenging in Class-Imbalanced Multi-Label Image Classifi…

Multi-Label Image ClassificationFew-Shot Image ClassificationMulti-Label ClassificationObject Recognition

Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models

2024-04-18 · Shouwei Ruan, Yinpeng Dong, Hanqing Liu, Yao Huang 외

Vision-Language Pre-training (VLP) models like CLIP have achieved remarkable success in computer vision and particularly demonstrated superior robustness to distribution shifts of 2D images. However, their robustness und…

GarmentZoom: Generating Zoomable Images from Garment Listings

2026-06-28 · Renjie Zhao, Jingwei Ma, Huy Huynh Cao, Brian Curless 외 arxiv

Online product listings for garments often include an overview photo and a close-up to show garment details. However, each photo focuses on either field of view or garment detail, forcing users to alternate between views…

Reference-based Super-Resolution