paper-with-me

Papers

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

2025-11-21 · Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang, Kai Zhao, Jing Xiao, Dan Zeng arxiv

Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, relative positions, and initial motion states. This so-called static regulation fails to capture dynamic changes of the object motion, including velocity changes and motion direction shifts. This limitation not only causes a temporal discrepancy between static references and dynamic vision modality but also constrains multi-modal tracking performance. To address this limitation, we propose a novel Vision-Motion-Reference aligned RMOT framework, named VMRMOT. It integrates a motion modality extracted from object dynamics to enhance the alignment between vision modality and language references through multi-modal large language models (MLLMs). Specifically, we introduce motion-aware descriptions derived from object dynamic behaviors and, leveraging the powerful temporal-reasoning capabilities of MLLMs, extract motion features as the motion modality. We further design a Vision-Motion-Reference Alignment (VMRA) module to hierarchically align visual queries with motion and reference cues, enhancing their cross-modal consistency. In addition, a Motion-Guided Prediction Head (MGPH) is developed to explore motion modality to enhance the performance of the prediction head. To the best of our knowledge, VMRMOT is the first approach to employ MLLMs in the RMOT task for vision-reference alignment. Extensive experiments on multiple RMOT benchmarks demonstrate that VMRMOT outperforms existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2511.17681

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Object Tracking

Similar Papers 제목 키워드 기반

Online Self-Preferring Language Models

2024-05-23 · Yuanzhao Zhai, Zhuo Zhang, Kele Xu, Hanyang Peng 외

Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward model to provide feedback for on-policy …

Using Lexical Alignment and Referring Ability to Address Data Sparsity in Situated Dialog Reference Resolution

2018-10-01 · EMNLP 2018 10 · Todd Shore, Gabriel Skantze

Referring to entities in situated dialog is a collaborative process, whereby interlocutors often expand, repair and/or replace referring expressions in an iterative process, converging on conceptual pacts of referring la…

Referring Expression

Mitigating Query Selection Bias in Referring Video Object Segmentation

2025-09-17 · Dingwei Zhang, Dong Zhang, Jinhui Tang arxiv

Recently, query-based methods have achieved remarkable performance in Referring Video Object Segmentation (RVOS) by using textual static object queries to drive cross-modal alignment. However, these static queries are ea…

Referring Video Object Segmentation

Leveraging Past References for Robust Language Grounding

2019-11-01 · CONLL 2019 11 · Subhro Roy, Michael Noseworthy, Rohan Paul, Daehyung Park 외

Grounding referring expressions to objects in an environment has traditionally been considered a one-off, ahistorical task. However, in realistic applications of grounding, multiple users will repeatedly refer to the sam…

ObjectReferring ExpressionVisual Grounding

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision

2025-01-08 · Giorgio Giannone, Ruoteng Li, Qianli Feng, Evgeny Perevodchikov 외

Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text trainin…

HallucinationQuestion AnsweringVisual Question Answering