Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation
How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion information uniformly. Specifically, AMC-Net fuses robust information from multi-modality features and promotes their collaboration in two stages. First, we propose a Multi-Modality Co-Attention Gate (MCG) on the bilateral encoder branches, in which a gate function is used to formulate co-attention scores for balancing the contributions of multi-modality features and suppressing the redundant and misleading information. Then, we propose a Motion Correction Module (MCM) with a visual-motion attention mechanism, which is constructed to emphasize the features of foreground objects by incorporating the spatio-temporal correspondence between appearance and motion cues. Extensive experiments on three public challenging benchmark datasets verify that our proposed network performs favorably against existing state-of-the-art methods via training with fewer data.
Code (1)
Tasks
Semantic SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationZero-Shot Video Object SegmentationSimilar Papers 제목 키워드 기반
AnimateZero: Video Diffusion Models are Zero-Shot Image Animators
Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes…
Image AnimationVideo GenerationMotion-Attentive Transition for Zero-Shot Video Object Segmentation
In this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object repres…
DecoderObjectSegmentationSemantic Segmentation+4Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation
Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in vide…
DenoisingPositionVideo GenerationAppearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model
Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body…
Zero-shot GeneralizationVideo ClassificationAction RecognitionZero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspe…
Optical Flow EstimationVideo EditingVideo Restoration