Weakly Supervised Video Salient Object Detection via Point Supervision
Video salient object detection models trained on pixel-wise dense annotation have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Several works attempt to use scribble annotations to mitigate this problem, but point supervision as a more labor-saving annotation method (even the most labor-saving method among manual annotation methods for dense prediction), has not been explored. In this paper, we propose a strong baseline model based on point supervision. To infer saliency maps with temporal information, we mine inter-frame complementary information from short-term and long-term perspectives, respectively. Specifically, we propose a hybrid token attention module, which mixes optical flow and image information from orthogonal directions, adaptively highlighting critical optical flow information (channel dimension) and critical token information (spatial dimension). To exploit long-term cues, we develop the Long-term Cross-Frame Attention module (LCFA), which assists the current frame in inferring salient objects based on multi-frame tokens. Furthermore, we label two point-supervised datasets, P-DAVIS and P-DAVSOD, by relabeling the DAVIS and the DAVSOD dataset. Experiments on the six benchmark datasets illustrate our method outperforms the previous state-of-the-art weakly supervised methods and even is comparable with some fully supervised approaches. Source code and datasets are available.
Code (0)
등록된 구현이 없습니다.
Tasks
Objectobject-detectionObject DetectionOptical Flow EstimationSalient Object DetectionVideo Salient Object DetectionSimilar Papers 제목 키워드 기반
Weakly Supervised Video Salient Object Detection
Significant performance improvement has been achieved for fully-supervised video salient object detection with the pixel-wise labeled training datasets, which are time-consuming and expensive to obtain. To relieve the bu…
Objectobject-detectionObject DetectionPseudo Label+4Encoding Based Saliency Detection for Videos and Images
We present a novel video saliency detection method to support human activity recognition and weakly supervised training of activity detection algorithms. Recent research has emphasized the need for analyzing salient inf…
Action DetectionActivity DetectionActivity RecognitionHuman Activity Recognition+6Learning Video Salient Object Detection Progressively from Unlabeled Videos
Recent deep learning-based video salient object detection (VSOD) has achieved some breakthrough, but these methods rely on expensive annotated videos with pixel-wise annotations, weak annotations, or part of the pixel-wi…
Objectobject-detectionObject DetectionOptical Flow Estimation+2Weakly Supervised Learning for Salient Object Detection
Recent advances in supervised salient object detection has resulted in significant performance on benchmark datasets. Training such models, however, requires expensive pixel-wise annotations of salient objects. Moreover,…
Objectobject-detectionObject DetectionRGB Salient Object Detection+3Weakly-Supervised Salient Object Detection Using Point Supervision
Current state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and labor-intensive. There are some weakly supervised m…
Objectobject-detectionObject DetectionSaliency Detection+1