paper-with-me

Papers

Temporal Attention for Cross-View Sequential Image Localization

2024-08-28 · Dong Yuan, Frederic Maire, Feras Dayoub

This paper introduces a novel approach to enhancing cross-view localization, focusing on the fine-grained, sequential localization of street-view images within a single known satellite image patch, a significant departure from traditional one-to-one image retrieval methods. By expanding to sequential image fine-grained localization, our model, equipped with a novel Temporal Attention Module (TAM), leverages contextual information to significantly improve sequential image localization accuracy. Our method shows substantial reductions in both mean and median localization errors on the Cross-View Image Sequence (CVIS) dataset, outperforming current state-of-the-art single-image localization techniques. Additionally, by adapting the KITTI-CVL dataset into sequential image sets, we not only offer a more realistic dataset for future research but also demonstrate our model's robust generalization capabilities across varying times and areas, evidenced by a 75.3% reduction in mean distance error in cross-view sequential image localization.

📄 PDF Abstract BibTeX arXiv:2408.15569

Code (1)

UQ-DongYuan/CVSeqLocation 공식 구현 pytorch

Tasks

Image RetrievalRetrieval

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Cross-View Image Sequence Geo-localization

2022-10-25 · Xiaohan Zhang, Waqas Sultani, Safwan Wshah

Cross-view geo-localization aims to estimate the GPS location of a query ground-view image by matching it to images from a reference database of geo-tagged aerial images. To address this challenging problem, recent appro…

geo-localization

End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration

2025-12-26 · Zhenwei Yang, Yibo Ai, Weidong Zhang arxiv

Multiview cooperative perception and multimodal fusion are essential for reliable 3-D spatiotemporal understanding in autonomous driving, especially in cases with occlusions, limited viewpoints, and communication delays …

Autonomous Driving

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

2025-07-23 · Chang Nie, Guangming Wang, Zhe Lie, Hesheng Wang arxiv

Robot imitation learning relies on 4D multi-view sequential images. However, the high cost of data collection and the scarcity of high-quality data severely constrain the generalization and application of embodied intell…

Data Augmentation

Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks

2017-12-12 · Bo Wu, Wen-Huang Cheng, Yongdong Zhang, Qiushi Huang 외

Prediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems. Previous research, although achieves pro…

PredictionSocial Media Popularity Prediction

RADIFUSION: A multi-radiomics deep learning based breast cancer risk prediction model using sequential mammographic images with image attention and bilateral asymmetry refinement

2023-04-01 · Hong Hui Yeoh, Andrea Liew, Raphaël Phan, Fredrik Strand 외

Breast cancer is a significant public health concern and early detection is critical for triaging high risk patients. Sequential screening mammograms can provide important spatiotemporal information about changes in brea…