paper-with-me

Papers

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

2026-04-29 · Xiaoya Cheng, Rouwan Wu, Xinyi Liu, Zeyu Cui, Yan Liu, Na Zhao, Yu Liu, Maojun Zhang, Shen Yan arxiv

Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale, high-fidelity training data. Existing benchmarks, predominantly biased toward ground-level or object-centric views, do not account for complex viewpoint transformations and diverse environmental conditions in UAV-based sensing. To bridge this critical gap, we propose AirZoo, a unified large-scale dataset and benchmark for grounding aerial geometric 3D vision. AirZoo possesses three appealing properties: 1) Scalable Generation Pipeline: Leveraging freely available, world-scale photogrammetric 3D meshes, it renders vast outdoor environments with customizable UAV flight trajectories and configurable weather/illumination. 2) Comprehensive Scene Diversity: It provides the most extensive coverage of region types to date (spanning 378 regions across 22 countries), systematically encompassing both highly structured urban landscapes and complex unstructured natural environments. 3) Rich Geometric Annotations: Each frame provides synchronized, pixel-level metric depth and precise 6-DoF geo-referenced poses, essential for geometry-aware learning. Through three rigorous evaluation tracks -- aerial image retrieval, cross-view matching, and multi-view 3D reconstruction -- we demonstrate that AirZoo serves as a powerful pre-training engine. Extensive experiments on both public and newly collected real-world benchmarks reveal that fine-tuning on AirZoo yields substantial performance gains for SoTA models (e.g., MegaLoc, RoMa, VGGT, and Depth Anything 3), establishing a new performance upper bound for aerial spatial intelligence.

📄 PDF Abstract BibTeX arXiv:2604.26567

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-View 3D ReconstructionImage Retrieval

Similar Papers 제목 키워드 기반

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

2026-04-09 · Joungbin An, Agrim Jain, Kristen Grauman arxiv

Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to overcome this limitation have adapted large multimodal language mod…

Video Grounding

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

2026-05-26 · Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei 외 arxiv

Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into multiple 1D tokens that are learned and decoded largely independently…

Visual Grounding

UniVTG: Towards Unified Video-Language Temporal Grounding

2023-07-31 · ICCV 2023 1 · Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick 외

Video Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing o…

Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1

Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset

2025-08-14 · Ziye Deng, Ruihan He, Jiaxiang Liu, Yuan Wang 외 arxiv

Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report gene…

Visual Question Answering

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

2024-11-21 · Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng 외

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encod…

Long-tailed Object DetectionObjectobject-detectionObject Detection+3