paper-with-me

Papers

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

2025-10-15 · Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh arxiv

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: 1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; 2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; 3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; 4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding.

📄 PDF Abstract BibTeX arXiv:2510.13243

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationSemantic SegmentationScene UnderstandingDomain Adaptation

Similar Papers 제목 키워드 기반

UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation

2024-04-22 · Siru Zhong, Xixuan Hao, Yibo Yan, Ying Zhang 외

Urbanization challenges underscore the necessity for effective satellite image-text retrieval methods to swiftly access specific information enriched with geographic semantics for urban applications. However, existing me…

DiversityDomain AdaptationImage-text RetrievalRetrieval+1

On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning Applications

2023-12-23 · Chenjiao Tan, Qian Cao, Yiwei Li, Jielu Zhang 외

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geogra…

geo-localizationimage-classificationImage ClassificationLand Cover Classification+5

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

2026-06-14 · Yanxin Xi, Xiang Su, Jie Feng, Yu Liu 외 arxiv

Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWe…

UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios

2024-08-30 · Baichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye 외

Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments. Moreover, existing urban benchmarks have bee…

Attributegeo-localizationScene Understanding

Deep Learning for Cross-Domain Data Fusion in Urban Computing: Taxonomy, Advances, and Outlook

2024-02-29 · Xingchen Zou, Yibo Yan, Xixuan Hao, Yuehong Hu 외

As cities continue to burgeon, Urban Computing emerges as a pivotal discipline for sustainable development by harnessing the power of cross-domain data fusion from diverse sources (e.g., geographical, traffic, social med…

Deep Learning