Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
Recent mainstream unsupervised video object segmentation (UVOS) motion-appearance approaches use either the bi-encoder structure to separately encode motion and appearance features, or the uni-encoder structure for joint encoding. However, these methods fail to properly balance the motion-appearance relationship. Consequently, even with complex fusion modules for motion-appearance integration, the extracted suboptimal features degrade the models' overall performance. Moreover, the quality of optical flow varies across scenarios, making it insufficient to rely solely on optical flow to achieve high-quality segmentation results. To address these challenges, we propose the Saliency-Motion guided Trunk-Collateral Network (SMTC-Net), which better balances the motion-appearance relationship and incorporates model's intrinsic saliency information to enhance segmentation performance. Specifically, considering that optical flow maps are derived from RGB images, they share both commonalities and differences. Accordingly, we propose a novel Trunk-Collateral structure for motion-appearance UVOS. The shared trunk backbone captures the motion-appearance commonality, while the collateral branch learns the uniqueness of motion features. Furthermore, an Intrinsic Saliency guided Refinement Module (ISRM) is devised to efficiently leverage the model's intrinsic saliency information to refine high-level features, and provide pixel-level guidance for motion-appearance fusion, thereby enhancing performance without additional input. Experimental results show that SMTC-Net achieved state-of-the-art performance on three UVOS datasets ( 89.2% J&F on DAVIS-16, 76% J on YouTube-Objects, 86.4% J on FBMS ) and four standard video salient object detection (VSOD) benchmarks with the notable increase, demonstrating its effectiveness and superiority over previous methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Flow EstimationSalient Object DetectionSemantic SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Salient Object DetectionVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
Unsupervised Video Object Segmentation using Motion Saliency-Guided Spatio-Temporal Propagation
Unsupervised video segmentation plays an important role in a wide variety of applications from object identification to compression. However, to date, fast motion, motion blur and occlusions pose significant challenges. …
Deep LearningOptical Flow EstimationSaliency PredictionSegmentation+6Unsupervised motion saliency map estimation based on optical flow inpainting
The paper addresses the problem of motion saliency in videos, that is, identifying regions that undergo motion departing from its context. We propose a new unsupervised paradigm to compute motion saliency maps. The key i…
Optical Flow EstimationRethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
This paper presents a fresh perspective on the role of saliency maps in weakly-supervised semantic segmentation (WSSS) and offers new insights and research directions based on our empirical findings. We conduct comprehen…
object-detectionObject DetectionSalient Object DetectionSemantic Segmentation+2DeFi Liquidation Risk Modeling Using the Reflection Principle for Zero-Drift Brownian Motion
In this paper, we propose an analytical method to compute the collateral liquidation probability in decentralized finance (DeFi) stablecoin single-collateral lending. Our approach models the collateral exchange rate as a…
Saliency-guided Emotion Modeling: Predicting Viewer Reactions from Video Stimuli
Understanding the emotional impact of videos is crucial for applications in content creation, advertising, and Human-Computer Interaction (HCI). Traditional affective computing methods rely on self-reported emotions, fac…