Motion Segmentation using Frequency Domain Transformer Networks
Self-supervised prediction is a powerful mechanism to learn representations that capture the underlying structure of the data. Despite recent progress, the self-supervised video prediction task is still challenging. One of the critical factors that make the task hard is motion segmentation, which is segmenting individual objects and the background and estimating their motion separately. In video prediction, the shape, appearance, and transformation of each object should be understood only by predicting the next frame in pixel space. To address this task, we propose a novel end-to-end learnable architecture that predicts the next frame by modeling foreground and background separately while simultaneously estimating and predicting the foreground motion using Frequency Domain Transformer Networks. Experimental evaluations show that this yields interpretable representations and that our approach can outperform some widely used video prediction methods like Video Ladder Network and Predictive Gated Pyramids on synthetic data.
Code (1)
Tasks
Motion SegmentationPredictionVideo PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Speech Emotion Recognition via an Attentive Time-Frequency Neural Network
Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time-frequency pattern of speech signal for speech emotion recognition (SER). \textcolor{black}{Generally, different e…
Emotion RecognitionSpeech Emotion RecognitionDomain Influence in MRI Medical Image Segmentation: spatial versus k-space inputs
Transformer-based networks applied to image patches have achieved cutting-edge performance in many vision tasks. However, lacking the built-in bias of convolutional neural networks (CNN) for local image statistics, they …
Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1FreqU-FNet: Frequency-Aware U-Net for Imbalanced Medical Image Segmentation
Medical image segmentation faces persistent challenges due to severe class imbalance and the frequency-specific distribution of anatomical structures. Most conventional CNN-based methods operate in the spatial domain and…
DecoderImage SegmentationMedical Image SegmentationSegmentation+1Generative adversarial network for segmentation of motion affected neonatal brain MRI
Automatic neonatal brain tissue segmentation in preterm born infants is a prerequisite for evaluation of brain development. However, automatic segmentation is often hampered by motion artifacts caused by infant head move…
Generative Adversarial NetworkImage ReconstructionImage SegmentationSegmentation+1FE-UNet: Frequency Domain Enhanced U-Net with Segment Anything Capability for Versatile Image Segmentation
Image segmentation is a critical task in visual understanding. Convolutional Neural Networks (CNNs) are predisposed to capture high-frequency features in images, while Transformers exhibit a contrasting focus on low-freq…
Image SegmentationSegmentationSemantic Segmentation