ASQuery: A Query-based Model for Action Segmentation
For the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem. In this paper, we propose a straight but effective model namely ASQuery by learning central representation of each action category, which transforms the classification problem to the similarity calculation between category-specific queries and frame features. These central representations are dynamically generated through our Transformer decoder module, endowing them more flexible and comprehensive perception of the whole video. Moreover, we first introduce the boundary query for refining segmentation results, aiding to alleviating the troublesome over-segmentation problem. ASQuery demonstrates superior performance compared to state-of-the-art models, achieving improvements of 0.9% and 4.1% in the mean metrics on two public action segmentation datasets, i.e., Breakfast and Assembly101, respectively. The source codes are available at https://github.com/zlngan/ASQuery.
Code (1)
Tasks
Action SegmentationDecodermodelSegmentationTemporal Action SegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Few-Shot Segmentation with Global and Local Contrastive Learning
In this work, we address the challenging task of few-shot segmentation. Previous few-shot segmentation methods mainly employ the information of support images as guidance for query image segmentation. Although some works…
Contrastive LearningImage SegmentationSegmentationSemantic SegmentationAsymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query
Actor and action video segmentation from natural language query aims to selectively segment the actor and its action in a video based on an input textual description. Previous works mostly focus on learning simple correl…
Referring Expression SegmentationSegmentationVideo SegmentationVideo Semantic SegmentationDynamic Prototype Convolution Network for Few-Shot Semantic Segmentation
The key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among support and query features and/or their prototypes, under the episodic training scenario. Most existing FSS method…
Few-Shot Semantic SegmentationSemantic SegmentationActor and Action Modular Network for Text-based Video Segmentation
Text-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitly align the video content with the textu…
Action SegmentationAction UnderstandingReferring Expression SegmentationSegmentation+3A Unified Query-based Paradigm for Camouflaged Instance Segmentation
Due to the high similarity between camouflaged instances and the background, the recently proposed camouflaged instance segmentation (CIS) faces challenges in accurate localization and instance segmentation. To this end,…
Boundary DetectionDecoderInstance SegmentationMulti-Task Learning+2