Learning Quintuplet Loss for Large-scale Visual Geo-Localization
With the maturity of Artificial Intelligence (AI) technology, Large Scale Visual Geo-Localization (LSVGL) is increasingly important in urban computing, where the task is to accurately and efficiently recognize the geo-location of a given query image. The main challenge of LSVGL faced by many experiments due to the appearance of real-word places may differ in various ways. While perspective deviation almost inevitably exists between training images and query images because of the arbitrary perspective. To cope with this situation, in this paper, we in-depth analyze the limitation of triplet loss which is the most commonly used metric learning loss in state-of-the-art LSVGL framework, and propose a new QUInTuplet Loss (QUITLoss) by embedding all the potential positive samples to the primitive triplet loss. Extensive experiments have been conducted to verify the effectiveness of the proposed approach and the results demonstrate that our new loss can enhance various LSVGL methods.
Code (0)
등록된 구현이 없습니다.
Tasks
geo-localizationMetric LearningTripletVisual Place RecognitionSimilar Papers 제목 키워드 기반
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when t…
Learning Joint Gait Representation via Quintuplet Loss Minimization
Gait recognition is an important biometric method popularly used in video surveillance, where the task is to identify people at a distance by their walking patterns from video sequences. Most of the current successful ap…
Gait RecognitionAdversarial Fine-Grained Composition Learning for Unseen Attribute-Object Recognition
Recognizing unseen attribute-object pairs never appearing in the training data is a challenging task, since an object often refers to a specific entity while an attribute is an abstract semantic description. Besides, att…
AttributeObjectObject RecognitionMulti-Level Visual Similarity Based Personalized Tourist Attraction Recommendation Using Geo-Tagged Photos
Geo-tagged photo based tourist attraction recommendation can discover users' travel preferences from their taken photos, so as to recommend suitable tourist attractions to them. However, existing visual content based met…
Towards Precise Intra-camera Supervised Person Re-identification
Intra-camera supervision (ICS) for person re-identification (Re-ID) assumes that identity labels are independently annotated within each camera view and no inter-camera identity association is labeled. It is a new settin…
Person Re-Identification