paper-with-me

홈 › Papers

TubeFormer-DeepLab: Video Mask Transformer

2022-05-30 · CVPR 2022 1 · Dahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao, Qihang Yu, Hong-Seok Kim, Hartwig Adam, In So Kweon, Liang-Chieh Chen

We present TubeFormer-DeepLab, the first attempt to tackle multiple core video segmentation tasks in a unified manner. Different video segmentation tasks (e.g., video semantic/instance/panoptic segmentation) are usually considered as distinct problems. State-of-the-art models adopted in the separate communities have diverged, and radically different approaches dominate in each task. By contrast, we make a crucial observation that video segmentation tasks could be generally formulated as the problem of assigning different predicted labels to video tubes (where a tube is obtained by linking segmentation masks along the time axis) and the labels may encode different values depending on the target task. The observation motivates us to develop TubeFormer-DeepLab, a simple and effective video mask transformer model that is widely applicable to multiple video segmentation tasks. TubeFormer-DeepLab directly predicts video tubes with task-specific labels (either pure semantic categories, or both semantic categories and instance identities), which not only significantly simplifies video segmentation models, but also advances state-of-the-art results on multiple video segmentation benchmarks

📄 PDF Abstract BibTeX arXiv:2205.15361

Code (0)

등록된 구현이 없습니다.

Tasks

Panoptic SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers

2020-12-01 · CVPR 2021 1 · Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille 외

We present MaX-DeepLab, the first end-to-end model for panoptic segmentation. Our approach simplifies the current pipeline that depends heavily on surrogate sub-tasks and hand-designed components, such as box detection, …

Panoptic Segmentation

CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation

2022-06-17 · CVPR 2022 1 · Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao 외

We propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer architectures used in segmentation and detect…

ClusteringPanoptic SegmentationSegmentation

kMaX-DeepLab: k-means Mask Transformer

2022-07-08 · Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins 외

The rise of transformers in vision tasks not only advances network backbone designs, but also starts a brand-new page to achieve end-to-end image recognition (e.g., object detection and panoptic segmentation). Originated…

ClusteringObject DetectionPanoptic SegmentationSemantic Segmentation

MaskConver: Revisiting Pure Convolution Model for Panoptic Segmentation

2023-12-11 · Abdullah Rashwan, Jiageng Zhang, Ali Taalimi, Fan Yang 외

In recent years, transformer-based models have dominated panoptic segmentation, thanks to their strong modeling capabilities and their unified representation for both semantic and instance classes as global binary masks.…

DecodermodelPanoptic Segmentation

Mask Propagation Network for Video Object Segmentation

2018-10-24 · Jia Sun, Dongdong Yu, Yinghong Li, Changhu Wang

In this work, we propose a mask propagation network to treat the video segmentation problem as a concept of the guided instance segmentation. Similar to most MaskTrack based video segmentation methods, our method takes t…

Instance SegmentationObjectOptical Flow EstimationSegmentation+4