Temporal Relational Modeling with Self-Supervision for Action Segmentation
Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) have shown promising advantages in relation reasoning on many tasks, it is still a challenge to apply graph convolution networks on long video sequences effectively. The main reason is that large number of nodes (i.e., video frames) makes GCNs hard to capture and model temporal relations in videos. To tackle this problem, in this paper, we introduce an effective GCN module, Dilated Temporal Graph Reasoning Module (DTGRM), designed to model temporal relations and dependencies between video frames at various time spans. In particular, we capture and model temporal relations via constructing multi-level dilated temporal graphs where the nodes represent frames from different moments in video. Moreover, to enhance temporal reasoning ability of the proposed model, an auxiliary self-supervised task is proposed to encourage the dilated temporal graph reasoning module to find and correct wrong temporal relations in videos. Our DTGRM model outperforms state-of-the-art action segmentation models on three challenging datasets: 50Salads, Georgia Tech Egocentric Activities (GTEA), and the Breakfast dataset. The code is available at https://github.com/redwang/DTGRM.
Code (1)
Tasks
Action RecognitionAction SegmentationAction UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling …
Contrastive LearningSkeleton Based Action SegmentationTAGTemporal Transformer Networks with Self-Supervision for Action Recognition
In recent years, 2D Convolutional Networks-based video action recognition has encouragingly gained wide popularity; However, constrained by the lack of long-range non-linear temporal relation modeling and reverse motion …
Action RecognitionTemporal Action LocalizationLearning Self-Similarity in Space and Time as Generalized Motion for Video Action Recognition
Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion r…
Action RecognitionTemporal Action LocalizationVideo UnderstandingRelational Temporal Graph Reasoning for Dual-task Dialogue Language Understanding
Dual-task dialog language understanding aims to tackle two correlative dialog language understanding tasks simultaneously via leveraging their inherent correlations. In this paper, we put forward a new framework, whose c…
Sentiment AnalysisSentiment ClassificationLearning Self-Similarity in Space and Time as a Generalized Motion for Action Recognition
Spatio-temporal convolution often fails to learn motion dynamics in videos and thus an effective motion representation is required for video understanding in the wild. In this paper, we propose a rich and robust motion r…
Action RecognitionVideo Understanding