paper-with-me

홈 › Papers

SeqFormer: Sequential Transformer for Video Instance Segmentation

2021-12-15 · Junfeng Wu, Yi Jiang, Song Bai, Wenqing Zhang, Xiang Bai

In this work, we present SeqFormer for video instance segmentation. SeqFormer follows the principle of vision transformer that models instance relationships among video frames. Nevertheless, we observe that a stand-alone instance query suffices for capturing a time sequence of instances in a video, but attention mechanisms shall be done with each frame independently. To achieve this, SeqFormer locates an instance in each frame and aggregates temporal information to learn a powerful representation of a video-level instance, which is used to predict the mask sequences on each frame dynamically. Instance tracking is achieved naturally without tracking branches or post-processing. On YouTube-VIS, SeqFormer achieves 47.4 AP with a ResNet-50 backbone and 49.0 AP with a ResNet-101 backbone without bells and whistles. Such achievement significantly exceeds the previous state-of-the-art performance by 4.6 and 4.4, respectively. In addition, integrated with the recently-proposed Swin transformer, SeqFormer achieves a much higher AP of 59.3. We hope SeqFormer could be a strong baseline that fosters future research in video instance segmentation, and in the meantime, advances this field with a more robust, accurate, neat model. The code is available at https://github.com/wjf5203/SeqFormer.

📄 PDF Abstract BibTeX arXiv:2112.08275

Code (2)

wjf5203/SeqFormer 공식 구현 pytorch
wjf5203/vnext pytorch

Tasks

Instance SegmentationSemantic SegmentationVideo Instance Segmentation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

Video Instance Segmentation in an Open-World

2023-04-03 · Omkar Thawakar, Sanath Narayan, Hisham Cholakkal, Rao Muhammad Anwer 외

Existing video instance segmentation (VIS) approaches generally follow a closed-world assumption, where only seen category instances are identified and spatio-temporally segmented at inference. Open-world formulation rel…

Instance SegmentationSemantic SegmentationVideo Instance Segmentation

Activity Graph Transformer for Temporal Action Localization

2021-01-21 · Megha Nawhal, Greg Mori

We introduce Activity Graph Transformer, an end-to-end learnable model for temporal action localization, that receives a video as input and directly predicts a set of action instances that appear in the video. Detecting …

Action LocalizationTemporal Action Localization

End-to-End Video Instance Segmentation with Transformers

2020-11-30 · CVPR 2021 1 · Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen 외

Video instance segmentation (VIS) is the task that requires simultaneously classifying, segmenting and tracking object instances of interest in video. Recent methods typically develop sophisticated pipelines to tackle th…

Instance SegmentationSegmentationSemantic SegmentationVideo Instance Segmentation+1

Robust Online Video Instance Segmentation with Track Queries

2022-11-16 · Zitong Zhan, Daniel McKee, Svetlana Lazebnik

Recently, transformer-based methods have achieved impressive results on Video Instance Segmentation (VIS). However, most of these top-performing methods run in an offline manner by processing the entire video clip at onc…

Image SegmentationInstance SegmentationMulti-Object TrackingObject Tracking+5

Sequential Clique Optimization for Video Object Segmentation

2018-09-01 · ECCV 2018 9 · Yeong Jun Koh, Young-Yoon Lee, Chang-Su Kim

A novel algorithm to segment out objects in a video sequence is proposed in this work. First, we extract object instances in each frame. Then, we select a visually important object instance in each frame to construct th…

Objectobject-detectionObject DetectionRGB Salient Object Detection+6