paper-with-me

Papers

Cross-Modal Object Tracking: Modality-Aware Representations and A Unified Benchmark

2021-11-08 · Chenglong Li, Tianhao Zhu, Lei Liu, Xiaonan Si, Zilin Fan, Sulan Zhai

In many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and infrared data is an effective way to handle imaging limitations of individual sources, but multi-modal imaging platforms usually require elaborate designs and cannot be applied in many real-world applications at present. Near-infrared (NIR) imaging becomes an essential part of many surveillance cameras, whose imaging is switchable between RGB and NIR based on the light intensity. These two modalities are heterogeneous with very different visual properties and thus bring big challenges for visual tracking. However, existing works have not studied this challenging problem. In this work, we address the cross-modal object tracking problem and contribute a new video dataset, including 654 cross-modal image sequences with over 481K frames in total, and the average video length is more than 735 frames. To promote the research and development of cross-modal object tracking, we propose a new algorithm, which learns the modality-aware target representation to mitigate the appearance gap between RGB and NIR modalities in the tracking process. It is plug-and-play and could thus be flexibly embedded into different tracking frameworks. Extensive experiments on the dataset are conducted, and we demonstrate the effectiveness of the proposed algorithm in two representative tracking frameworks against 17 state-of-the-art tracking methods. We will release the dataset for free academic usage, dataset download link and code will be released soon.

📄 PDF Abstract BibTeX arXiv:2111.04264

Code (0)

등록된 구현이 없습니다.

Tasks

Object TrackingVisual Tracking

Similar Papers 제목 키워드 기반

Cross-Modal Object Tracking via Modality-Aware Fusion Network and A Large-Scale Dataset

2023-12-22 · Lei Liu, Mengya Zhang, Cheng Li, Chenglong Li 외

Visual tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrar…

Object TrackingVisual Tracking

Learning Dual-Fused Modality-Aware Representations for RGBD Tracking

2022-11-06 · Shang Gao, Jinyu Yang, Zhe Li, Feng Zheng 외

With the development of depth sensors in recent years, RGBD object tracking has received significant attention. Compared with the traditional RGB object tracking, the addition of the depth modality can effectively solve …

Object Tracking

Multi-Adapter RGBT Tracking

2019-07-17 · Chenglong Li, Andong Lu, Aihua Zheng, Zhengzheng Tu 외

The task of RGBT tracking aims to take the complementary advantages from visible spectrum and thermal infrared data to achieve robust visual tracking, and receives more and more attention in recent years. Existing works …

Visual Tracking

Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking

2026-03-10 · Shilei Wang, Pujian Lai, Dong Gao, Jifeng Ning 외 arxiv

Most existing multimodal trackers adopt uniform fusion strategies, overlooking the inherent differences between modalities. Moreover, they propagate temporal information through mixed tokens, leading to entangled and les…

Object Tracking

Tracking and Segmenting Anything in Any Modality

2025-11-22 · Tianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong Han arxiv

Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approache…

Representation LearningMulti-Object Tracking