paper-with-me

홈 › Papers

WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory

2025-09-16 · Ruifei Ding, Zhe Chen, Wen Fan, Chen Long, Huijuan Xiao, Yelu Zeng, Zhen Dong, Bisheng Yang arxiv

Street trees are vital to urban livability, providing ecological and social benefits. Establishing a detailed, accurate, and dynamically updated street tree inventory has become essential for optimizing these multifunctional assets within space-constrained urban environments. Given that traditional field surveys are time-consuming and labor-intensive, automated surveys utilizing Mobile Mapping Systems (MMS) offer a more efficient solution. However, existing MMS-acquired tree datasets are limited by small-scale scene, limited annotation, or single modality, restricting their utility for comprehensive analysis. To address these limitations, we introduce WHU-STree, a cross-city, richly annotated, and multi-modal urban street tree dataset. Collected across two distinct cities, WHU-STree integrates synchronized point clouds and high-resolution images, encompassing 21,007 annotated tree instances across 50 species and 2 morphological parameters. Leveraging the unique characteristics, WHU-STree concurrently supports over 10 tasks related to street tree inventory. We benchmark representative baselines for two key tasks--tree species classification and individual tree segmentation. Extensive experiments and in-depth analysis demonstrate the significant potential of multi-modal data fusion and underscore cross-domain applicability as a critical prerequisite for practical algorithm deployment. In particular, we identify key challenges and outline potential future works for fully exploiting WHU-STree, encompassing multi-modal fusion, multi-task collaboration, cross-domain generalization, spatial pattern learning, and Multi-modal Large Language Model for street tree asset management. The WHU-STree dataset is accessible at: https://github.com/WHU-USI3DV/WHU-STree.

📄 PDF Abstract BibTeX arXiv:2509.13172

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationPoint Clouds

Similar Papers 제목 키워드 기반

StreetReaderAI: Making Street View Accessible Using Context-Aware Multimodal AI

2025-08-11 · Jon E. Froehlich, Alexander Fiannaca, Nimer Jaber, Victor Tsaran 외 arxiv

Interactive streetscape mapping tools such as Google Street View (GSV) and Meta Mapillary enable users to virtually navigate and experience real-world environments via immersive 360° imagery but remain fundamentally inac…

Designing streetscapes from street-view imagery using diffusion models

2026-05-17 · Yuzhou Chen, Yuebing Liang, Lingqian Hu, Kailai Sun 외 arxiv

Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existing studies largely focus on measuring current streetscapes and rarel…

Scene Generation

MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark

2025-05-18 · Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao 외

Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, lack multimodal diversity and underrepresent dense, mixed-use street-level spaces, especially in non-Western urban contexts.…

Multimodal ReasoningVisual Place Recognition

Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

2026-05-25 · Niels Sombekke, Rob G. J. Wijnhoven, Martin R. Oswald arxiv

We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial patch tokens from a shared DINOv2 backbone. The design naturally ha…

Multi-modal Classification

m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning

2026-01-27 · Yosub Shin, Michael Buriek, Igor Molybog arxiv

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We intr…

Reinforcement LearningSpatial Reasoning