paper-with-me

홈 › Papers

Mamba-based Spatio-Frequency Motion Perception for Video Camouflaged Object Detection

2025-07-31 · Xin Li, Keren Fu, Qijun Zhao arxiv

Existing video camouflaged object detection (VCOD) methods primarily rely on spatial appearances for motion perception. However, the high foreground-background similarity in VCOD limits the discriminability of such features (e.g. color and texture). Recent studies demonstrate that frequency features can not only compensate for appearance limitations, but also perceive motion through dynamic variations in spectral energy. Meanwhile, the emerging state space model called Mamba enables efficient motion perception in frame sequences with its linear-time long-sequence modeling capability. Motivated by this, we propose Vcamba, a visual camouflage Mamba based on spatio-frequency motion perception that integrates frequency and spatial features for efficient and accurate VCOD. Specifically, by analyzing the spatial representations of frequency components, we reveal a structural evolution pattern that emerges from the ordered superposition of components. Based on this observation, we propose a unique frequency-domain sequential scanning (FSS) strategy to unfold the spectrum. Utilizing FSS, the adaptive frequency enhancement (AFE) module employs Mamba to model the causal dependencies within sequences, enabling effective frequency learning. Furthermore, we propose a space-based long-range motion perception (SLMP) module and a frequency-based long-range motion perception (FLMP) module to model spatio-temporal and frequency-temporal sequences. Finally, the space and frequency motion fusion module (SFMF) integrates dual-domain features into unified motion representation. Experiments show that Vcamba outperforms state-of-the-art methods across 6 evaluation metrics on 2 datasets with lower computation cost, confirming its superiority. Code is available at: https://github.com/BoydeLi/Vcamba.

📄 PDF Abstract BibTeX arXiv:2507.23601

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal SequencesObject Detection

Similar Papers 제목 키워드 기반

FTDMamba: Frequency-Assisted Temporal Dilation Mamba for Unmanned Aerial Vehicle Video Anomaly Detection

2026-01-16 · Cheng-Zhuang Liu, Si-Bao Chen, Qing-Ling Shu, Chris Ding 외 arxiv

Recent advances in video anomaly detection (VAD) mainly focus on ground-based surveillance or unmanned aerial vehicle (UAV) videos with static backgrounds, whereas research on UAV videos with dynamic backgrounds remains …

Video Anomaly Detection

High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation

2025-10-13 · Runyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim 외 arxiv

Modeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential …

Pose Estimation

DemMamba: Alignment-free Raw Video Demoireing with Frequency-assisted Spatio-Temporal Mamba

2024-08-20 · Shuning Xu, Xina Liu, Binbin Song, Xiangyu Chen 외

Moire patterns, resulting from the interference of two similar repetitive patterns, are frequently observed during the capture of images or videos on screens. These patterns vary in color, shape, and location across vide…

MambaOptical Flow Estimation

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

2025-11-09 · Hua Chang, Xin Xu, Wei Liu, Wei Wang 외 arxiv

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hinderi…

Space-time Video Super-resolution

MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition

2025-10-12 · Deng Li, Jun Shao, Bohao Xing, Rong Gao 외 arxiv

Micro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal dependencies. While CNNs are effective at cap…

Micro-gesture Recognition