TAEC: Unsupervised Action Segmentation with Temporal-Aware Embedding and Clustering
Temporal action segmentation in untrimmed videos has gained increased attention recently. However, annotating action classes and frame-wise boundaries is extremely time consuming and cost intensive, especially on large-scale datasets. To address this issue, we propose an unsupervised approach for learning action classes from untrimmed video sequences. In particular, we propose a temporal embedding network that combines relative time prediction, feature reconstruction, and sequence-to-sequence learning, to preserve the spatial layout and sequential nature of the video features. A two-step clustering pipeline on these embedded feature representations then allows us to enforce temporal consistency within, as well as across videos. Based on the identified clusters, we decode the video into coherent temporal segments that correspond to semantically meaningful action classes. Our evaluation on three challenging datasets shows the impact of each component and, furthermore, demonstrates our state-of-the-art unsupervised action segmentation results.
Code (0)
등록된 구현이 없습니다.
Tasks
Action SegmentationClusteringSegmentationTemporal Action SegmentationUnsupervised Action SegmentationSimilar Papers 제목 키워드 기반
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
This paper presents an unsupervised transformer-based framework for temporal activity segmentation which leverages not only frame-level cues but also segment-level cues. This is in contrast with previous methods which of…
Action SegmentationDecoderPredictionSegmentation+1Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
Current state-of-the-art methods for skeleton-based temporal action segmentation are predominantly supervised and require annotated data, which is expensive to collect. In contrast, existing unsupervised temporal action …
Action SegmentationTemporally Consistent Unbalanced Optimal Transport for Unsupervised Action Segmentation
We propose a novel approach to the action segmentation task for long, untrimmed videos, based on solving an optimal transport problem. By encoding a temporal consistency prior into a Gromov-Wasserstein problem, we are ab…
Action SegmentationSegmentationUnsupervised Action SegmentationSMC-NCA: Semantic-guided Multi-level Contrast for Semi-supervised Temporal Action Segmentation
Semi-supervised temporal action segmentation (SS-TA) aims to perform frame-wise classification in long untrimmed videos, where only a fraction of videos in the training set have labels. Recent studies have shown the pote…
Action SegmentationContrastive LearningRepresentation LearningSegmentation+1Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels…
Action Segmentation