paper-with-me

홈 › Papers

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

2026-03-20 · Jingyu Guo, Ziye Chen, Ziwen Li, Zhengqing Gao, Jiaxin Huang, Hanlue Zhang, Fengming Huang, Yu Yao, Tongliang Liu, Mingming Gong arxiv

Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descriptions with goal-centric evaluation, making them less diagnostic for real operations where brief, high-level commands must be grounded into safe multi-stage behaviors. We present HUGE-Bench, a benchmark for High-Level UAV Vision-Language-Action (HL-VLA) tasks that tests whether an agent can interpret concise language and execute complex, process-oriented trajectories with safety awareness. HUGE-Bench comprises 4 real-world digital twin scenes, 8 high-level tasks, and 2.56M meters of trajectories, and is built on an aligned 3D Gaussian Splatting (3DGS)-Mesh representation that combines photorealistic rendering with collision-capable geometry for scalable generation and collision-aware evaluation. We introduce process-oriented and collision-aware metrics to assess process fidelity, terminal accuracy, and safety. Experiments on representative state-of-the-art VLA models reveal significant gaps in high-level semantic completion and safe execution, highlighting HUGE-Bench as a diagnostic testbed for high-level UAV autonomy.

📄 PDF Abstract BibTeX arXiv:2603.19822

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language Navigation

Similar Papers 제목 키워드 기반

SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer

2025-02-11 · Wenxi Li, Yuchen Guo, Jilai Zheng, Haozhe Lin 외

Recent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shots in the MS COCO dataset, the higher res…

object-detectionObject Detection

SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

2023-06-26 · NeurIPS 2023 11 · Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi 외

In the last year alone, a surge of new benchmarks to measure compositional understanding of vision-language models have permeated the machine learning ecosystem. Given an image, these benchmarks probe a model's ability t…

TransNAS-Bench-101: Improving Transferrability and Generalizability of Cross-Task Neural Architecture Search

2021-01-01 · Yawen Duan, Xin Chen, Hang Xu, Zewei Chen 외

Recent breakthroughs of Neural Architecture Search (NAS) are extending the field's research scope towards a broader range of vision tasks and more diversified search spaces. While existing NAS methods mostly design archi…

GPUNeural Architecture SearchTransfer Learning

OmniMAE: Single Model Masked Pretraining on Images and Videos

2022-06-16 · CVPR 2023 1 · Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh, Kalyan Vasudev Alwala 외

Transformer-based architectures have become competitive across a variety of visual domains, most notably images and videos. While prior work studies these modalities in isolation, having a common architecture suggests th…

MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding

2024-01-01 · CVPR 2024 1 · Xu Cao, Tong Zhou, Yunsheng Ma, Wenqian Ye 외

Vision-language generative AI has demonstrated remarkable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However current benchmark datasets lack mul…

Autonomous DrivingInstruction FollowingPrompt EngineeringScene Understanding