Motion-aware Memory Network for Fast Video Salient Object Detection
Previous methods based on 3DCNN, convLSTM, or optical flow have achieved great success in video salient object detection (VSOD). However, they still suffer from high computational costs or poor quality of the generated saliency maps. To solve these problems, we design a space-time memory (STM)-based network, which extracts useful temporal information of the current frame from adjacent frames as the temporal branch of VSOD. Furthermore, previous methods only considered single-frame prediction without temporal association. As a result, the model may not focus on the temporal information sufficiently. Thus, we initially introduce object motion prediction between inter-frame into VSOD. Our model follows standard encoder--decoder architecture. In the encoding stage, we generate high-level temporal features by using high-level features from the current and its adjacent frames. This approach is more efficient than the optical flow-based methods. In the decoding stage, we propose an effective fusion strategy for spatial and temporal branches. The semantic information of the high-level features is used to fuse the object details in the low-level features, and then the spatiotemporal features are obtained step by step to reconstruct the saliency maps. Moreover, inspired by the boundary supervision commonly used in image salient object detection (ISOD), we design a motion-aware loss for predicting object boundary motion and simultaneously perform multitask learning for VSOD and object motion prediction, which can further facilitate the model to extract spatiotemporal features accurately and maintain the object integrity. Extensive experiments on several datasets demonstrated the effectiveness of our method and can achieve state-of-the-art metrics on some datasets. The proposed model does not require optical flow or other preprocessing, and can reach a speed of nearly 100 FPS during inference.
Code (1)
Tasks
motion predictionObjectobject-detectionObject DetectionOptical Flow EstimationSalient Object DetectionVideo Salient Object DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TransFlow: Motion Knowledge Transfer from Video Diffusion Models to Video Salient Object Detection
Video salient object detection (SOD) relies on motion cues to distinguish salient objects from backgrounds, but training such models is limited by scarce video datasets compared to abundant image datasets. Existing appro…
Video Salient Object DetectionRegion Aware Video Object Segmentation with Deep Motion Modeling
Current semi-supervised video object segmentation (VOS) methods usually leverage the entire features of one frame to predict object masks and update memory. This introduces significant redundant computations. To reduce r…
DecoderObjectSegmentationSemantic Segmentation+3Self-Supervised Video Object Segmentation by Motion-Aware Mask Propagation
We propose a self-supervised spatio-temporal matching method, coined Motion-Aware Mask Propagation (MAMP), for video object segmentation. MAMP leverages the frame reconstruction task for training without the need for ann…
SegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1Motion-Appearance Co-Memory Networks for Video Question Answering
Video Question Answering (QA) is an important task in understanding video temporal structure. We observe that there are three unique attributes of video QA compared with image QA: (1) it deals with long sequences of imag…
Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)Motion Guided Attention for Video Salient Object Detection
Video salient object detection aims at discovering the most visually distinctive objects in a video. How to effectively take object motion into consideration during video salient object detection is a critical issue. Exi…
Objectobject-detectionObject DetectionOptical Flow Estimation+4