FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in videos. We formulate this task as a structured prediction problem and design a two-stream fully convolutional neural network which fuses together motion and appearance in a unified framework. Since large-scale video datasets with pixel level segmentations are problematic, we show how to bootstrap weakly annotated videos together with existing image recognition datasets for training. Through experiments on three challenging video segmentation benchmarks, our method substantially improves the state-of-the-art for segmenting generic (unseen) objects. Code and pre-trained models are available on the project website.
Code (0)
등록된 구현이 없습니다.
Tasks
SegmentationStructured PredictionUnsupervised Video Object SegmentationVideo SegmentationVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
FusionSeg: Learning to Combine Motion and Appearance for Fully Automatic Segmentation of Generic Objects in Videos
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in …
SegmentationStructured PredictionVideo SegmentationVideo Semantic SegmentationFusionSegReID: Advancing Person Re-Identification with Multimodal Retrieval and Precise Segmentation
Person re-identification (ReID) plays a critical role in applications like security surveillance and criminal investigations by matching individuals across large image galleries captured by non-overlapping cameras. Tradi…
Person Re-IdentificationRetrievalGeneral Automatic Human Shape and Motion Capture Using Volumetric Contour Cues
Markerless motion capture algorithms require a 3D body with properly personalized skeleton dimension and/or body shape and appearance to successfully track a person. Unfortunately, many tracking methods consider model pe…
Markerless Motion CaptureAdaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip Reading
The goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical f…
Landmark-based LipreadingLip ReadingOptical Flow EstimationAutomatic Face Reenactment
We propose an image-based, facial reenactment system that replaces the face of an actor in an existing target video with the face of a user from a source video, while preserving the original target performance. Our syste…
ClusteringFace ModelFace ReenactmentFace Transfer+2