From Two-Stream to One-Stream: Efficient RGB-T Tracking via Mutual Prompt Learning and Knowledge Distillation
Due to the complementary nature of visible light and thermal infrared modalities, object tracking based on the fusion of visible light images and thermal images (referred to as RGB-T tracking) has received increasing attention from researchers in recent years. How to achieve more comprehensive fusion of information from the two modalities at a lower cost has been an issue that researchers have been exploring. Inspired by visual prompt learning, we designed a novel two-stream RGB-T tracking architecture based on cross-modal mutual prompt learning, and used this model as a teacher to guide a one-stream student model for rapid learning through knowledge distillation techniques. Extensive experiments have shown that, compared to similar RGB-T trackers, our designed teacher model achieved the highest precision rate, while the student model, with comparable precision rate to the teacher model, realized an inference speed more than three times faster than the teacher model.(Codes will be available if accepted.)
Code (0)
등록된 구현이 없습니다.
Tasks
Knowledge DistillationObject TrackingPrompt LearningRgb-T TrackingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
RGB-T Tracking via Multi-Modal Mutual Prompt Learning
Object tracking based on the fusion of visible and thermal im-ages, known as RGB-T tracking, has gained increasing atten-tion from researchers in recent years. How to achieve a more comprehensive fusion of information fr…
Object TrackingPrompt LearningRgb-T TrackingVisual Prompt Multi-Modal Tracking
Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is f…
Object TrackingPrompt LearningRgb-T TrackingMiddle Fusion and Multi-Stage, Multi-Form Prompts for Robust RGB-T Tracking
RGB-T tracking, a vital downstream task of object tracking, has made remarkable progress in recent years. Yet, it remains hindered by two major challenges: 1) the trade-off between performance and efficiency; 2) the scar…
FormObject TrackingPrompt LearningRgb-T TrackingVL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feature extraction and rely heavily on impli…
Visual TrackingFine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-Identification
Recently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is ch…
Cloth-Changing Person Re-IdentificationPerson Re-Identification