paper-with-me

Papers

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

2026-08-21 · Yutian Jiang, Jiabo Liu, Xixuan Hao, Yuxuan Liang arxiv

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.

📄 PDF Abstract BibTeX arXiv:2608.21041

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

2025-10-15 · Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh arxiv

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting…

Monocular Depth EstimationSemantic SegmentationScene UnderstandingDomain Adaptation

Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery

2025-11-14 · Yijie Kang, Xinliang Wang, Zhenyu Wu, Yifeng Shi 외 arxiv

3D urban generation from satellite imagery is an important task for scalable digital twins and real-world simulation environments. Existing approaches primarily rely on scene-level generation paradigms, which often requi…

Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges

2020-09-07 · CVPR 2021 1 · Qingyong Hu, Bo Yang, Sheikh Khalid, Wen Xiao 외

An essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly avai…

Scene UnderstandingSemantic Segmentation

UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models

2025-05-16 · Yuhang Liu, Yingxue Zhang, Xin Zhang, Ling Tian 외

Understanding and predicting urban dynamics is crucial for managing transportation systems, optimizing urban planning, and enhancing public services. While neural network-based approaches have achieved success, they ofte…

Test-time Adaptation

3D Scene Understanding at Urban Intersection using Stereo Vision and Digital Map

2021-12-10 · Prarthana Bhattacharyya, Yanlei Gu, Jiali Bao, Xu Liu 외

The driving behavior at urban intersections is very complex. It is thus crucial for autonomous vehicles to comprehensively understand challenging urban traffic scenes in order to navigate intersections and prevent accide…

Autonomous VehiclesNavigateScene Understanding