paper-with-me

홈 › Papers

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

2026-05-18 · Quan Zhang, Zeqiang Cai, Peiming Zhao, Jingze Wu, Cailun Wu, Hongbo Chen, Jianhuang Lai arxiv

Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across views to achieve robustness. However, view-invariant inherently enforces part-level alignment, which ignores view-specific cues and discriminative identity information. To this end, this work proposes ViSA (View-aware Semantic Alignment), a view-aware framework that achieves cross-view semantic consistency containing an Expert-driven Token Generation Module (ETGM) and a Dual-branch Local Fusion Module (DLFM). Technically, the former constructs a set of view-aware experts to generate adaptive semantic queries that perceive viewpoint-specific patterns, while the latter leverages graph reasoning to extract and align local regions responsive to different experts. Extensive experiments on three AGPReID benchmarks including AG-ReID.v2, CARGO and LAGPeR demonstrate that ViSA consistently achieves superior performance, with a notable 10.06\% mAP improvement on the challenging CARGO cross-view protocol. The code is available at \href{https://github.com/Cat-Zero/ViSA}{https://github.com/Cat-Zero/ViSA}.

📄 PDF Abstract BibTeX arXiv:2605.18192

Code (0)

등록된 구현이 없습니다.

Tasks

Person Re-Identification

Similar Papers 제목 키워드 기반

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

2025-10-25 · Qiao Li, Jie Li, Yukang Zhang, Lei Tan 외 arxiv

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and…

Person Re-Identification

Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark

2026-03-21 · Yifei Deng, Chenglong Li, Yuyang Zhang, Guyue Hu 외 arxiv

Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. Compared to ground-view text--image perso…

Person RetrievalText Generation

SMDT: Cross-View Geo-Localization with Image Alignment and Transformer

2022-04-06 · IEEE International Conference on Multimedia and Expo 2022 2022 4 · Xiaoyang Tian, Jie Shao, Deqiang Ouyang, Anjie Zhu 외

The goal of cross-view geo-localization is to determine the location of a given ground image by matching with aerial images. However, existing methods ignore the variability of scenes, additional information and spatial …

geo-localizationSegmentationSemantic Segmentation

Semantic-aware Network for Aerial-to-Ground Image Synthesis

2023-08-14 · Jinhyun Jang, Taeyong Song, Kwanghoon Sohn

Aerial-to-ground image synthesis is an emerging and challenging problem that aims to synthesize a ground image from an aerial image. Due to the highly different layout and object representation between the aerial and gro…

Image Generation

Top2Ground: A Height-Aware Dual Conditioning Diffusion Model for Robust Aerial-to-Ground View Generation

2025-11-11 · Jae Joong Lee, Bedrich Benes arxiv

Generating ground-level images from aerial views is a challenging task due to extreme viewpoint disparity, occlusions, and a limited field of view. We introduce Top2Ground, a novel diffusion-based method that directly ge…