paper-with-me

홈 › Papers

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models

2026-04-02 · Qiyao Zhang, Shuhua Zheng, Jianli Sun, Chengxiang Li, Xianke Wu, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian arxiv

Embodied visual tracking is crucial for Unmanned Aerial Vehicles (UAVs) executing complex real-world tasks. In dynamic urban scenarios with complex semantic requirements, Vision-Language-Action (VLA) models show great promise due to their cross-modal fusion and continuous action generation capabilities. To benchmark multimodal tracking in such environments, we construct a dedicated evaluation benchmark and a large-scale dataset encompassing over 890K frames, 176 tasks, and 85 diverse objects. Furthermore, to address temporal feature redundancy and the lack of spatial geometric priors in existing VLA models, we propose an improved VLA tracking model, UAV-Track VLA. Built upon the $π_{0.5}$ architecture, our model introduces a temporal compression net to efficiently capture inter-frame dynamics. Additionally, a parallel dual-branch decoder comprising a spatial-aware auxiliary grounding head and a flow matching action expert is designed to decouple cross-modal features and generate fine-grained continuous actions. Systematic experiments in the CARLA simulator validate the superior end-to-end performance of our method. Notably, in challenging long-distance pedestrian tracking tasks, UAV-Track VLA achieves a 61.76\% success rate and 269.65 average tracking frames, significantly outperforming existing baselines. Furthermore, it demonstrates robust zero-shot generalization in unseen environments and reduces single-step inference latency by 33.4\% (to 0.0571s) compared to the original $π_{0.5}$, enabling highly efficient, real-time UAV control. Data samples and demonstration videos are available at: https://github.com/Hub-Tian/UAV-Track_VLA.

📄 PDF Abstract BibTeX arXiv:2604.02241

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationVisual Tracking

Similar Papers 제목 키워드 기반

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

2025-11-26 · Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun 외 arxiv

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current …

Representation LearningMulti-Object TrackingObject Detection

DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking

2026-05-17 · Guyue Hu, Haoming Liu, Siyuan Song, Chenglong Li 외 arxiv

Aerial object tracking has broad applications in public safety, emergency rescue, wildlife monitoring, and related fields. However, existing aerial tracking benchmarks are mainly based on passive 2D video sequences captu…

Object Tracking

TrackVLA: Embodied Visual Tracking in the Wild

2025-05-29 · Shaoan Wang, Jiazhao Zhang, Minghan Li, Jiahang Liu 외

Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inherently challenging as it requires both …

Language ModelingLanguage ModellingObject RecognitionTrajectory Planning+2

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

2025-05-27 · Kui Wu, Shuhang Xu, Hao Chen, Churan Wang 외

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracki…

Spatial ReasoningVisual Tracking

Unsupervised Domain Adaptation for Nighttime Aerial Tracking

2022-03-20 · CVPR 2022 1 · Junjie Ye, Changhong Fu, Guangze Zheng, Danda Pani Paudel 외

Previous advances in object tracking mostly reported on favorable illumination circumstances while neglecting performance at nighttime, which significantly impeded the development of related aerial robot applications. Th…

Domain AdaptationObject DiscoveryObject TrackingUnsupervised Domain Adaptation