paper-with-me

홈 › Papers

AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios

2025-11-26 · Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun, Haocheng Zhao, Haiyun Jiang, Tao Huang, Henghui Ding, Qing-Long Han arxiv

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrains their ability to capture broad-scale scene contexts and perform comprehensive tracking and path planning. In contrast, Unmanned Aerial Vehicles (UAVs) leverage their expansive aerial perspectives and superior maneuverability to enable wide-area surveillance. Moreover, UAVs have emerged as critical platforms for Embodied Intelligence, which has given rise to an unprecedented demand for intelligent aerial systems capable of natural language interaction. To this end, we introduce AerialMind, the first large-scale RMOT benchmark in UAV scenarios, which aims to bridge this research gap. To facilitate its construction, we develop an innovative semi-automated collaborative agent-based labeling assistant (COALA) framework that significantly reduces labor costs while maintaining annotation quality. Furthermore, we propose HawkEyeTrack (HETrack), a novel method that collaboratively enhances vision-language representation learning and improves the perception of UAV scenarios. Comprehensive experiments validated the challenging nature of our dataset and the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2511.21053

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningMulti-Object TrackingObject Detection

Similar Papers 제목 키워드 기반

STORM: End-to-End Referring Multi-Object Tracking in Videos

2026-04-12 · Zijia Lu, Jingru Yi, Jue Wang, Yuxiao Chen 외 arxiv

Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding …

Multi-Object Tracking

Cross-View Referring Multi-Object Tracking

2024-12-23 · Sijia Chen, En Yu, Wenbing Tao

Referring Multi-Object Tracking (RMOT) is an important topic in the current tracking field. Its task form is to guide the tracker to track objects that match the language description. Current research mainly focuses on r…

Cross-view Referring Multi-Object TrackingMulti-Object TrackingObjectObject Tracking+1

Bootstrapping Referring Multi-Object Tracking

2024-06-07 · Yani Zhang, Dongming Wu, Wencheng Han, Xingping Dong

Referring multi-object tracking (RMOT) aims at detecting and tracking multiple objects following human instruction represented by a natural language expression. Existing RMOT benchmarks are usually formulated through man…

DiversityMulti-Object TrackingObjectObject Tracking+1

RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking

2026-02-25 · Yanqiu Yu, Zhifan Jin, Sijia Chen, Tongfei Chu 외 arxiv

Referring Multi-Object Tracking has attracted increasing attention due to its human-friendly interactive characteristics, yet it exhibits limitations in low-visibility conditions, such as nighttime, smoke, and other chal…

Multi-Object Tracking

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

2025-12-11 · Henghui Ding, Chang Liu, Shuting He, Kaining Ying 외 arxiv

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Ex…

Referring Video Object SegmentationMulti-Object TrackingVideo SegmentationVideo Captioning