paper-with-me

홈 › Papers

Middle Fusion and Multi-Stage, Multi-Form Prompts for Robust RGB-T Tracking

2024-03-27 · Qiming Wang, Yongqiang Bai, Hongxing Song

RGB-T tracking, a vital downstream task of object tracking, has made remarkable progress in recent years. Yet, it remains hindered by two major challenges: 1) the trade-off between performance and efficiency; 2) the scarcity of training data. To address the latter challenge, some recent methods employ prompts to fine-tune pre-trained RGB tracking models and leverage upstream knowledge in a parameter-efficient manner. However, these methods inadequately explore modality-independent patterns and disregard the dynamic reliability of different modalities in open scenarios. We propose M3PT, a novel RGB-T prompt tracking method that leverages middle fusion and multi-modal and multi-stage visual prompts to overcome these challenges. We pioneer the use of the adjustable middle fusion meta-framework for RGB-T tracking, which could help the tracker balance the performance with efficiency, to meet various demands of application. Furthermore, based on the meta-framework, we utilize multiple flexible prompt strategies to adapt the pre-trained model to comprehensive exploration of uni-modal patterns and improved modeling of fusion-modal features in diverse modality-priority scenarios, harnessing the potential of prompt learning in RGB-T tracking. Evaluating on 6 existing challenging benchmarks, our method surpasses previous state-of-the-art prompt fine-tuning methods while maintaining great competitiveness against excellent full-parameter fine-tuning methods, with only 0.34M fine-tuned parameters.

📄 PDF Abstract BibTeX arXiv:2403.18193

Code (0)

등록된 구현이 없습니다.

Tasks

FormObject TrackingPrompt LearningRgb-T Tracking

Similar Papers 제목 키워드 기반

RTMaps-based Local Dynamic Map for multi-ADAS data fusion

2022-05-13 · Marcos Nieto, Mikel Garcia, Itziar Urbieta, Oihana Otaegui

Work on Local Dynamic Maps (LDM) implementation is still in its early stages, as the LDM standards only define how information shall be structured in databases, while the mechanism to fuse or link information across diff…

Decision Making

Middle-level Fusion for Lightweight RGB-D Salient Object Detection

2021-04-23 · Nianchang Huang, Qiang Zhang, Jungong Han

Most existing lightweight RGB-D salient object detection (SOD) models are based on two-stream structure or single-stream structure. The former one first uses two sub-networks to extract unimodal features from RGB and dep…

object-detectionObject DetectionRGB-D Salient Object DetectionSalient Object Detection

CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection

2026-06-28 · Yu Li, Yuenan Hou, Yingmei Wei, Jiangming Chen 외 arxiv

Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of the conventional RGB frames. However, previous multi-modal fusion techniques typi…

Object Detection

GFD-SSD: Gated Fusion Double SSD for Multispectral Pedestrian Detection

2019-03-16 · Yang Zheng, Izzat H. Izzat, Shahrzad Ziaee

Pedestrian detection is an essential task in autonomous driving research. In addition to typical color images, thermal images benefit the detection in dark environments. Hence, it is worthwhile to explore an integrated a…

Autonomous DrivingPedestrian Detection

Depth Super-Resolution from Explicit and Implicit High-Frequency Features

2023-03-16 · Xin Qiao, Chenyang Ge, Youmin Zhang, Yanhui Zhou 외

We propose a novel multi-stage depth super-resolution network, which progressively reconstructs high-resolution depth maps from explicit and implicit high-frequency features. The former are extracted by an efficient tran…

Super-ResolutionVocal Bursts Intensity Prediction