paper-with-me

홈 › Papers

MambaEVT: Event Stream based Visual Object Tracking using State Space Model

2024-08-20 · Xiao Wang, Chao Wang, Shiao Wang, Xixi Wang, Zhicheng Zhao, Lin Zhu, Bo Jiang

Event camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Current event-based tracking algorithms are gradually hitting their performance bottlenecks, due to the utilization of vision Transformer and the static template for target object localization. In this paper, we propose a novel Mamba-based visual tracking framework that adopts the state space model with linear complexity as a backbone network. The search regions and target template are fed into the vision Mamba network for simultaneous feature extraction and interaction. The output tokens of search regions will be fed into the tracking head for target localization. More importantly, we consider introducing a dynamic template update strategy into the tracking framework using the Memory Mamba network. By considering the diversity of samples in the target template library and making appropriate adjustments to the template memory module, a more effective dynamic template can be integrated. The effective combination of dynamic and static templates allows our Mamba-based tracking algorithm to achieve a good balance between accuracy and computational cost on multiple large-scale datasets, including EventVOT, VisEvent, and FE240hz. The source code will be released on https://github.com/Event-AHU/MambaEVT

📄 PDF Abstract BibTeX arXiv:2408.10487

Code (1)

event-ahu/mambaevt 공식 구현 pytorch

Tasks

MambaObject LocalizationObject TrackingVisual Object TrackingVisual Tracking

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Adversarial Attack for RGB-Event based Visual Object Tracking

2025-04-19 · Qiang Chen, Xiao Wang, Haowen Wang, Bo Jiang 외

Visual object tracking is a crucial research topic in the fields of computer vision and multi-modal fusion. Among various approaches, robust visual tracking that combines RGB frames with Event streams has attracted incre…

Adversarial AttackObject TrackingVisual Object TrackingVisual Tracking

Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

2024-03-09 · Xiao Wang, Ju Huang, Shiao Wang, Chuanming Tang 외

Current event-/frame-event based trackers undergo evaluation on short-term tracking datasets, however, the tracking of real-world scenarios involves long-term tracking, and the performance of existing tracking algorithms…

Object TrackingRgb-T TrackingVisual Tracking

Spatial Orthogonal Refinement for Robust RGB-Event Visual Object Tracking

2026-03-29 · Dexing Huang, Shiao Wang, Fan Zhang, Xiao Wang arxiv

Robust visual object tracking (VOT) remains challenging in high-speed motion scenarios, where conventional RGB sensors suffer from severe motion blur and performance degradation. Event cameras, with microsecond temporal …

Visual Object Tracking

Robust event-stream pattern tracking based on correlative filter

2018-03-17 · Hongmin Li, Luping Shi

Object tracking based on retina-inspired and event-based dynamic vision sensor (DVS) is challenging for the noise events, rapid change of event-stream shape, chaos of complex background textures, and occlusion. To addres…

Event-based visionObjectObject Tracking

Visual Prompt Multi-Modal Tracking

2023-03-20 · CVPR 2023 1 · Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang 외

Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is f…

Object TrackingPrompt LearningRgb-T Tracking