paper-with-me

홈 › Papers

LaMOT: Language-Guided Multi-Object Tracking

2024-06-12 · Yunhao Li, Xiaoqiong Liu, Luke Liu, Heng Fan, Libo Zhang

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional use of templates or pre-set information from training sets in conventional tracking tasks. Despite various efforts, a key challenge lies in the lack of a clear understanding of why language is used for tracking, which hinders further development in this field. In this paper, we address this challenge by introducing Language-Guided MOT, a unified task framework, along with a corresponding large-scale benchmark, termed LaMOT, which encompasses diverse scenarios and language descriptions. Specially, LaMOT comprises 1,660 sequences from 4 different datasets and aims to unify various Vision-Language MOT tasks while providing a standardized evaluation platform. To ensure high-quality annotations, we manually assign appropriate descriptive texts to each target in every video and conduct careful inspection and correction. To the best of our knowledge, LaMOT is the first benchmark dedicated to Language-Guided MOT. Additionally, we propose a simple yet effective tracker, termed LaMOTer. By establishing a unified task framework, providing challenging benchmarks, and offering insights for future algorithm design and evaluation, we expect to contribute to the advancement of research in Vision-Language MOT. We will release the data at https://github.com/Nathan-Li123/LaMOT.

📄 PDF Abstract BibTeX arXiv:2406.08324

Code (1)

nathan-li123/lamot 공식 구현

Tasks

DescriptiveMulti-Object TrackingObjectObject Tracking

Similar Papers 제목 키워드 기반

GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System

2024-08-17 · Shuo Wang, Yongcai Wang, Zhimin Xu, Yongyu Guo 외

For interacting with mobile objects in unfamiliar environments, simultaneously locating, mapping, and tracking the 3D poses of multiple objects are crucially required. This paper proposes a Tracklet Graph and Query Graph…

Multiple Object TrackingObjectobject-detectionObject Detection+1

VLAMotor: Test-Guided Enhancement of Vision-Language-Action Models via Agent-BasedData Synthesis

2026-05-16 · Zeqin Liao, Peifan Ren, Zixu Gao, Hongyu Gong 외 arxiv

Vision-Language-Action (VLA) models follow a data-driven paradigm and are constrained by the coverage of training data, making them prone to failure on edge-case configurations after deployment. To mitigate such risks, i…

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

2026-07-28 · Xiangqun Zhang, Likai Wang, Zekun Qian, Ruize Han 외 arxiv

Referring multi-object tracking (RMOT) extends tracking from category-driven perception to language-guided understanding by grounding object trajectories in natural-language expressions. Despite recent progress, existing…

Multi-Object TrackingObject Detection

VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement

2025-09-17 · Jun Du, Weiwei Xing, Ming Li, Fei Richard Yu arxiv

Current multi-object tracking (MOT) algorithms typically overlook issues inherent in low-quality videos, leading to significant degradation in tracking performance when confronted with real-world image deterioration. The…

Multi-Object Tracking

Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance

2026-08-19 · Kai Van Brunt, Justin Kay, Sara Beery arxiv

Computer vision is increasingly used to automate recognition tasks in large ecological datasets, but more complex tasks such as multi-object tracking continue to pose challenges. As researchers seek to incorporate vision…

Multi-Object Tracking