paper-with-me

홈 › Papers

Segmenting the motion components of a video: A long-term unsupervised model

2023-10-02 · Etienne Meunier, Patrick Bouthemy

Human beings have the ability to continuously analyze a video and immediately extract the motion components. We want to adopt this paradigm to provide a coherent and stable motion segmentation over the video sequence. In this perspective, we propose a novel long-term spatio-temporal model operating in a totally unsupervised way. It takes as input the volume of consecutive optical flow (OF) fields, and delivers a volume of segments of coherent motion over the video. More specifically, we have designed a transformer-based network, where we leverage a mathematically well-founded framework, the Evidence Lower Bound (ELBO), to derive the loss function. The loss function combines a flow reconstruction term involving spatio-temporal parametric motion models combining, in a novel way, polynomial (quadratic) motion models for the spatial dimensions and B-splines for the time dimension of the video sequence, and a regularization term enforcing temporal consistency on the segments. We report experiments on four VOS benchmarks, demonstrating competitive quantitative results, while performing motion segmentation on a whole sequence in one go. We also highlight through visual results the key contributions on temporal consistency brought by our method.

📄 PDF Abstract BibTeX arXiv:2310.01040

Code (0)

등록된 구현이 없습니다.

Tasks

Motion SegmentationOptical Flow EstimationRepresentation Learning

Methods 이 논문이 사용한 방법론

VOS VOS is a type of video object segmentation model consisting of two network components. The target appearance model consists of a light-weight module, which is learned during…

Similar Papers 제목 키워드 기반

Learning segmentation from point trajectories

2025-01-21 · Laurynas Karazija, Iro Laina, Christian Rupprecht, Andrea Vedaldi

We consider the problem of segmenting objects in videos based on their motion and no other forms of supervision. Prior work has often approached this problem by using the principle of common fate, namely the fact that th…

Optical Flow Estimation

Spatial Decomposition and Temporal Fusion based Inter Prediction for Learned Video Compression

2024-01-29 · Xihua Sheng, Li Li, Dong Liu, Houqiang Li

Video compression performance is closely related to the accuracy of inter prediction. It tends to be difficult to obtain accurate inter prediction for the local video regions with inconsistent motion and occlusion. Tradi…

Motion EstimationMS-SSIMPredictionSSIM+1

2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-20 · Bin Cao, Yisi Zhang, Xuanxu Lin, Xingjian He 외

Motion Expression guided Video Segmentation is a challenging task that aims at segmenting objects in the video based on natural language expressions with motion descriptions. Unlike the previous referring video object se…

Instance SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation+4

STFCN: Spatio-Temporal FCN for Semantic Video Segmentation

2016-08-21 · Mohsen Fayyaz, Mohammad Hajizadeh Saffar, Mohammad Sabokrou, Mahmood Fathy 외

This paper presents a novel method to involve both spatial and temporal features for semantic video segmentation. Current work on convolutional neural networks(CNNs) has shown that CNNs provide advanced spatial features …

SegmentationSemantic SegmentationVideo SegmentationVideo Semantic Segmentation

FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos

2017-01-19 · CVPR 2017 · Suyog Dutt Jain, Bo Xiong, Kristen Grauman

We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …

SegmentationStructured PredictionUnsupervised Video Object SegmentationVideo Segmentation+1