paper-with-me

Papers

Learning Street View Representations with Spatiotemporal Contrast

2025-02-07 · Yong Li, Yingjing Huang, Gengchen Mai, Fan Zhang

Street view imagery is extensively utilized in representation learning for urban visual environments, supporting various sustainable development tasks such as environmental perception and socio-economic assessment. However, it is challenging for existing image representations to specifically encode the dynamic urban environment (such as pedestrians, vehicles, and vegetation), the built environment (including buildings, roads, and urban infrastructure), and the environmental ambiance (such as the cultural and socioeconomic atmosphere) depicted in street view imagery to address downstream tasks related to the city. In this work, we propose an innovative self-supervised learning framework that leverages temporal and spatial attributes of street view imagery to learn image representations of the dynamic urban environment for diverse downstream tasks. By employing street view images captured at the same location over time and spatially nearby views at the same time, we construct contrastive learning tasks designed to learn the temporal-invariant characteristics of the built environment and the spatial-invariant neighborhood ambiance. Our approach significantly outperforms traditional supervised and unsupervised methods in tasks such as visual place recognition, socioeconomic estimation, and human-environment perception. Moreover, we demonstrate the varying behaviors of image representations learned through different contrastive learning objectives across various downstream tasks. This study systematically discusses representation learning strategies for urban studies based on street view images, providing a benchmark that enhances the applicability of visual data in urban science. The code is available at https://github.com/yonglleee/UrbanSTCL.

📄 PDF Abstract BibTeX arXiv:2502.04638

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningRepresentation LearningSelf-Supervised LearningVisual Place Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Diagnosing Urban Street Vitality via a Visual-Semantic and Spatiotemporal Framework for Street-Level Economics

2026-04-10 · Xinxin Zhuo, Mengyuan Niu, Ruizhe Wang, Junyan Yang 외 arxiv

Micro-scale street-level economic assessment is fundamental for precision spatial resource allocation. While Street View Imagery (SVI) advances urban sensing, existing approaches remain semantically superficial and overl…

Instance Segmentation

Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood Embedding

2020-01-29 · Zhecheng Wang, Haoyuan Li, Ram Rajagopal

Understanding intrinsic patterns and predicting spatiotemporal characteristics of cities require a comprehensive representation of urban neighborhoods. Existing works relied on either inter- or intra-region connectivitie…

Document EmbeddingSemantic SimilaritySemantic Textual Similarity

OpenStreetView-5M: The Many Roads to Global Visual Geolocation

2024-04-29 · CVPR 2024 1 · Guillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn 외

Determining the location of an image anywhere on Earth is a complex visual task, which makes it particularly relevant for evaluating computer vision algorithms. Yet, the absence of standard, large-scale, open-access data…

Photo geolocation estimation

Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior

2025-04-24 · Lin Che, Yizi Chen, Tanhua Jin, Martin Raubal 외

Urban land use classification and mapping are critical for urban planning, resource management, and environmental monitoring. Existing remote sensing techniques often lack precision in complex urban environments due to t…

Clustering

Graph representation learning for street networks

2022-11-09 · Mateo Neira, Roberto Murcio

Streets networks provide an invaluable source of information about the different temporal and spatial patterns emerging in our cities. These streets are often represented as graphs where intersections are modelled as nod…

DecoderGraph Representation LearningRepresentation Learning