Understanding Multi-Task Activities from Single-Task Videos
(MT-TAS), a novel paradigm that addresses the challenges of interleaved actions when performing multiple tasks simultaneously. Traditional action segmentation models, trained on single-task videos, struggle to handle task switches and complex scenes inherent in multi-task scenarios. To overcome these challenges, our MT-TAS approach synthesizes multi-task video data from single-task sources using our Multi-task Sequence Blending and Segment Boundary Learning modules. Additionally, we propose to dynamically isolate foreground and background elements within video frames, addressing the intricacies of object layouts in multi-task scenarios and enabling a new two-stage temporal action segmentation framework with Foreground-Aware Action Refinement. Also, we introduce the Multi-task Egocentric Kitchen Activities (MEKA) dataset, containing 12 hours of egocentric multi-task videos, to rigorously benchmark MT-TAS models. Extensive experiments demonstrate that our framework effectively bridges the gap between single-task training and multi-task testing, advancing temporal action segmentation with state-of-the-art performance in complex environments.
Code (0)
등록된 구현이 없습니다.
Tasks
Action SegmentationSegmentationTemporal Action SegmentationSimilar Papers 제목 키워드 기반
LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task Activities
Understanding and interpreting human actions is a long-standing challenge and a critical indicator of perception in artificial intelligence. However, a few imperative components of daily human activities are largely miss…
Action RecognitionAction UnderstandingHuman-Object Interaction DetectionLEMMA+1The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose
The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human activities, existing public datasets, whil…
Action RecognitionObjectPose EstimationSegmentation+2MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing
Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabi…
Activity RecognitionFew Shot Action RecognitionGraph GenerationVideo UnderstandingDiscovery of Shared Semantic Spaces for Multi-Scene Video Query and Summarization
The growing rate of public space CCTV installations has generated a need for automated methods for exploiting video surveillance data including scene understanding, query, behaviour annotation and summarization. For this…
Scene UnderstandingSemantic SimilaritySemantic Textual SimilarityVideo SummarizationUncertainty-Aware Anticipation of Activities
Anticipating future activities in video is a task with many practical applications. While earlier approaches are limited to just a few seconds in the future, the prediction time horizon has just recently been extended to…