paper-with-me

Papers

Understanding Multi-Task Activities from Single-Task Videos

2025-01-01 · CVPR 2025 1 · YuHan Shen, Ehsan Elhamifar

(MT-TAS), a novel paradigm that addresses the challenges of interleaved actions when performing multiple tasks simultaneously. Traditional action segmentation models, trained on single-task videos, struggle to handle task switches and complex scenes inherent in multi-task scenarios. To overcome these challenges, our MT-TAS approach synthesizes multi-task video data from single-task sources using our Multi-task Sequence Blending and Segment Boundary Learning modules. Additionally, we propose to dynamically isolate foreground and background elements within video frames, addressing the intricacies of object layouts in multi-task scenarios and enabling a new two-stage temporal action segmentation framework with Foreground-Aware Action Refinement. Also, we introduce the Multi-task Egocentric Kitchen Activities (MEKA) dataset, containing 12 hours of egocentric multi-task videos, to rigorously benchmark MT-TAS models. Extensive experiments demonstrate that our framework effectively bridges the gap between single-task training and multi-task testing, advancing temporal action segmentation with state-of-the-art performance in complex environments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action SegmentationSegmentationTemporal Action Segmentation

Similar Papers 제목 키워드 기반

LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task Activities

2020-07-31 · ECCV 2020 8 · Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu 외

Understanding and interpreting human actions is a long-standing challenge and a critical indicator of perception in artificial intelligence. However, a few imperative components of daily human activities are largely miss…

Action RecognitionAction UnderstandingHuman-Object Interaction DetectionLEMMA+1

The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

2020-07-01 · Yizhak Ben-Shabat, Xin Yu, Fatemeh Sadat Saleh, Dylan Campbell 외

The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human activities, existing public datasets, whil…

Action RecognitionObjectPose EstimationSegmentation+2

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

2022-11-28 · NeurIPS 2022 11 · Zelun Luo, Zane Durante, Linden Li, Wanze Xie 외

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabi…

Activity RecognitionFew Shot Action RecognitionGraph GenerationVideo Understanding

Discovery of Shared Semantic Spaces for Multi-Scene Video Query and Summarization

2015-07-27 · Xun Xu, Timothy Hospedales, Shaogang Gong

The growing rate of public space CCTV installations has generated a need for automated methods for exploiting video surveillance data including scene understanding, query, behaviour annotation and summarization. For this…

Scene UnderstandingSemantic SimilaritySemantic Textual SimilarityVideo Summarization

Uncertainty-Aware Anticipation of Activities

2019-08-26 · Yazan Abu Farha, Juergen Gall

Anticipating future activities in video is a task with many practical applications. While earlier approaches are limited to just a few seconds in the future, the prediction time horizon has just recently been extended to…