paper-with-me

Papers

Towards General Multimodal Visual Tracking

2025-03-14 · Andong Lu, Mai Wen, Jinhu Wang, Yuanzhi Guo, Chenglong Li, Jin Tang, Bin Luo

Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through leveraging complementary cues from different sources, it remains challenging in complex scenes due to the limitations of bi-modal scenarios. In this work, we introduce a general multimodal visual tracking task that fully exploits the advantages of four modalities, including RGB, thermal infrared, event, and language, for robust tracking under challenging conditions. To provide a comprehensive evaluation platform for general multimodal visual tracking, we construct QuadTrack600, a large-scale, high-quality benchmark comprising 600 video sequences (totaling 384.7K high-resolution (640x480) frame groups). In each frame group, all four modalities are spatially aligned and meticulously annotated with bounding boxes, while 21 sequence-level challenge attributes are provided for detailed performance analysis. Despite quad-modal data provides richer information, the differences in information quantity among modalities and the computational burden from four modalities are two challenging issues in fusing four modalities. To handle these issues, we propose a novel approach called QuadFusion, which incorporates an efficient Multiscale Fusion Mamba with four different scanning scales to achieve sufficient interactions of the four modalities while overcoming the exponential computational burden, for general multimodal visual tracking. Extensive experiments on the QuadTrack600 dataset and three bi-modal tracking datasets, including LasHeR, VisEvent, and TNL2K, validate the effectiveness of our QuadFusion.

📄 PDF Abstract BibTeX arXiv:2503.11218

Code (0)

등록된 구현이 없습니다.

Tasks

MambaVisual Tracking

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

DreamTrack: Dreaming the Future for Multimodal Visual Object Tracking

2025-01-01 · CVPR 2025 1 · Mingzhe Guo, Weiping Tan, Wenyu Ran, Liping Jing 외

Aiming to achieve class-agnostic perception in visual object tracking, current trackers commonly formulate tracking as a one-shot detection problem with the template-matching architecture. Despite the success, severe…

Object TrackingTemplate MatchingVisual Object TrackingVisual Tracking

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

2024-03-24 · CVPR 2024 1 · Xiaojun Hou, Jiazheng Xing, Yijie Qian, Yaowei Guo 외

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers, which was inefficient and lacked generalized repres…

Object TrackingRgb-T TrackingVisual Object Tracking

MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models

2025-02-15 · Vanya Cohen, Raymond Mooney

Entity tracking is a fundamental challenge in natural language understanding, requiring models to maintain coherent representations of entities. Previous work has benchmarked entity tracking performance in purely text-ba…

Natural Language UnderstandingVisual Reasoning

OneThinker: All-in-one Reasoning Model for Image and Video

2025-12-02 · Kaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan 외 arxiv

Reinforcement learning (RL) has recently achieved remarkable success in eliciting visual reasoning within Multimodal Large Language Models (MLLMs). However, existing approaches typically train separate models for differe…

Zero-shot GeneralizationReinforcement LearningMultimodal ReasoningQuestion Answering

Reliable Object Tracking by Multimodal Hybrid Feature Extraction and Transformer-Based Fusion

2024-05-28 · Hongze Sun, Rui Liu, Wuque Cai, Jun Wang 외

Visual object tracking, which is primarily based on visible light image sequences, encounters numerous challenges in complicated scenarios, such as low light conditions, high dynamic ranges, and background clutter. To ad…

ObjectObject TrackingVisual Object Tracking