paper-with-me

홈 › Papers

SkyScenes: A Synthetic Dataset for Aerial Scene Understanding

2023-12-11 · Sahil Khose, Anisha Pal, Aayushi Agarwal, Deepanshi, Judy Hoffman, Prithvijit Chattopadhyay

Real-world aerial scene understanding is limited by a lack of datasets that contain densely annotated images curated under a diverse set of conditions. Due to inherent challenges in obtaining such images in controlled real-world settings, we present SkyScenes, a synthetic dataset of densely annotated aerial images captured from Unmanned Aerial Vehicle (UAV) perspectives. We carefully curate SkyScenes images from CARLA to comprehensively capture diversity across layouts (urban and rural maps), weather conditions, times of day, pitch angles and altitudes with corresponding semantic, instance and depth annotations. Through our experiments using SkyScenes, we show that (1) models trained on SkyScenes generalize well to different real-world scenarios, (2) augmenting training on real images with SkyScenes data can improve real-world performance, (3) controlled variations in SkyScenes can offer insights into how models respond to changes in viewpoint conditions (height and pitch), weather and time of day, and (4) incorporating additional sensor modalities (depth) can improve aerial scene understanding. Our dataset and associated generation code are publicly available at: https://hoffman-group.github.io/SkyScenes/

📄 PDF Abstract BibTeX arXiv:2312.06719

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityScene Understanding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…
CARLA CARLA is an open-source simulator for autonomous driving research. CARLA has been developed from the ground up to support development, training, and validation of autonomous urban…

Similar Papers 제목 키워드 기반

Quantifying the synthetic and real domain gap in aerial scene understanding

2024-11-29 · Alina Marcu

Quantifying the gap between synthetic and real-world imagery is essential for improving both transformer-based models - that rely on large volumes of data - and datasets, especially in underexplored domains like aerial s…

Domain AdaptationScene Understanding

ClaraVid: A Holistic Scene Reconstruction Benchmark From Aerial Perspective With Delentropy-Based Complexity Profiling

2025-03-22 · Radu Beche, Sergiu Nedevschi

The development of aerial holistic scene understanding algorithms is hindered by the scarcity of comprehensive datasets that enable both semantic and geometric reconstruction. While synthetic datasets offer an alternativ…

Panoptic SegmentationScene Understanding

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

2025-10-15 · Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh arxiv

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting…

Monocular Depth EstimationSemantic SegmentationScene UnderstandingDomain Adaptation

AID: A Benchmark Dataset for Performance Evaluation of Aerial Scene Classification

2016-08-18 · Gui-Song Xia, Jingwen Hu, Fan Hu, Baoguang Shi 외

Aerial scene classification, which aims to automatically label an aerial image with a specific semantic category, is a fundamental problem for understanding high-resolution remote sensing imagery. In recent years, it has…

Aerial Scene ClassificationClassificationGeneral ClassificationScene Classification

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?

2025-04-25 · Yusen Zhang, Wenliang Zheng, Aashrith Madasu, Peng Shi 외

High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language…

Diagnostic