Global2Local: Efficient Structure Search for Video Action Segmentation
Temporal receptive fields of models play an important role in action segmentation. Large receptive fields facilitate the long-term relations among video clips while small receptive fields help capture the local details. Existing methods construct models with hand-designed receptive fields in layers. Can we effectively search for receptive field combinations to replace hand-designed patterns? To answer this question, we propose to find better receptive field combinations through a global-to-local search scheme. Our search scheme exploits both global search to find the coarse combinations and local search to get the refined receptive field combination patterns further. The global search finds possible coarse combinations other than human-designed patterns. On top of the global search, we propose an expectation guided iterative local search scheme to refine combinations effectively. Our global-to-local search can be plugged into existing action segmentation methods to achieve state-of-the-art performance.
Code (2)
Tasks
Action SegmentationSegmentationSimilar Papers 제목 키워드 기반
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community have been made towards a local-to-global pa…
Video CaptioningActBERT: Learning Global-Local Video-Text Representations
In this paper, we introduce ActBERT for self-supervised learning of joint video-text representations from unlabeled data. First, we leverage global action information to catalyze the mutual interactions between linguisti…
Action SegmentationQuestion AnsweringRepresentation LearningRetrieval+3Video Contrastive Learning with Global Context
Contrastive learning has revolutionized self-supervised image representation learning field, and recently been adapted to video domain. One of the greatest advantages of contrastive learning is that it allows us to flexi…
Action ClassificationAction LocalizationContrastive LearningRepresentation Learning+2T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions. The key to this problem is to measure text-video similarities in a joint embedding space. How…
RetrievalVideo RetrievalLeveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction
Predicting an interaction before it is fully executed is very important in applications such as human-robot interaction and video surveillance. In a two-human interaction scenario, there often contextual dependency struc…
Optical Flow Estimation