paper-with-me

홈 › Papers

Spatio-Temporal Multi-Task Learning Transformer for Joint Moving Object Detection and Segmentation

2021-06-21 · Eslam Mohamed, Ahmed El-Sallab

Moving objects have special importance for Autonomous Driving tasks. Detecting moving objects can be posed as Moving Object Segmentation, by segmenting the object pixels, or Moving Object Detection, by generating a bounding box for the moving targets. In this paper, we present a Multi-Task Learning architecture, based on Transformers, to jointly perform both tasks through one network. Due to the importance of the motion features to the task, the whole setup is based on a Spatio-Temporal aggregation. We evaluate the performance of the individual tasks architecture versus the MTL setup, both with early shared encoders, and late shared encoder-decoder transformers. For the latter, we present a novel joint tasks query decoder transformer, that enables us to have tasks dedicated heads out of the shared model. To evaluate our approach, we use the KITTI MOD [29] data set. Results show1.5% mAP improvement for Moving Object Detection, and 2%IoU improvement for Moving Object Segmentation, over the individual tasks networks.

📄 PDF Abstract BibTeX arXiv:2106.11401

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDecoderMoving Object DetectionMulti-Task LearningObjectobject-detectionObject DetectionSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

TubeDETR: Spatio-Temporal Video Grounding with Transformers

2022-03-30 · CVPR 2022 1 · Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 외

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal …

DecoderLanguage-Based Temporal LocalizationNatural Language Visual Groundingobject-detection+6

Spatio-Temporal Tuples Transformer for Skeleton-Based Action Recognition

2022-01-08 · Helei Qiu, Biao Hou, Bo Ren, Xiaohua Zhang

Capturing the dependencies between joints is critical in skeleton-based action recognition task. Transformer shows great potential to model the correlation of important joints. However, the existing Transformer-based met…

Action RecognitionSkeleton Based Action Recognition

MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video

2022-03-02 · CVPR 2022 1 · Jinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen 외

Recent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the m…

3D Human Pose EstimationClassificationMonocular 3D Human Pose EstimationPose Estimation

GLSFormer : Gated - Long, Short Sequence Transformer for Step Recognition in Surgical Videos

2023-07-20 · Nisarg A. Shah, Shameema Sikder, S. Swaroop Vedula, Vishal M. Patel

Automated surgical step recognition is an important task that can significantly improve patient safety and decision-making during surgeries. Existing state-of-the-art methods for surgical step recognition either rely on …

Decision Making

STAR-Transformer: A Spatio-temporal Cross Attention Transformer for Human Action Recognition

2022-10-14 · WACV 2023 1 · Dasom Ahn, Sangwon Kim, Hyunsu Hong, Byoung Chul Ko

In action recognition, although the combination of spatio-temporal videos and skeleton features can improve the recognition performance, a separate model and balancing feature representation for cross-modal data are requ…

Action RecognitionDecoderTemporal Action Localization