paper-with-me

홈 › Papers

Distill and Collect for Semi-Supervised Temporal Action Segmentation

2022-11-02 · Sovan Biswas, Anthony Rhodes, Ramesh Manuvinakurike, Giuseppe Raffa, Richard Beckwith

Recent temporal action segmentation approaches need frame annotations during training to be effective. These annotations are very expensive and time-consuming to obtain. This limits their performances when only limited annotated data is available. In contrast, we can easily collect a large corpus of in-domain unannotated videos by scavenging through the internet. Thus, this paper proposes an approach for the temporal action segmentation task that can simultaneously leverage knowledge from annotated and unannotated video sequences. Our approach uses multi-stream distillation that repeatedly refines and finally combines their frame predictions. Our model also predicts the action order, which is later used as a temporal constraint while estimating frames labels to counter the lack of supervision for unannotated videos. In the end, our evaluation of the proposed approach on two different datasets demonstrates its capability to achieve comparable performance to the full supervision despite limited annotation.

📄 PDF Abstract BibTeX arXiv:2211.01311

Code (0)

등록된 구현이 없습니다.

Tasks

Action SegmentationSegmentationTemporal Action Segmentation

Similar Papers 제목 키워드 기반

Learning from Temporal Gradient for Semi-supervised Action Recognition

2021-11-25 · CVPR 2022 1 · Junfei Xiao, Longlong Jing, Lin Zhang, Ju He 외

Semi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-bas…

Action RecognitionTemporal Action Localization

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

2023-03-28 · CVPR 2023 1 · Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah

Semi-Supervised Learning can be more beneficial for the video domain compared to images because of its higher annotation cost and dimensionality. Besides, any video understanding task requires reasoning over both spatial…

Action RecognitionOptical Flow EstimationVideo Understanding

Semi-Supervised Vision-Language-Action Model

2026-06-19 · Hongyang He, Jiuming Liu, Victor Sanchez arxiv

Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to new environments still depends on costly action-labeled demonstration…

Knowledge-Spreader: Learning Semi-Supervised Facial Action Dynamics by Consistifying Knowledge Granularity

2023-01-01 · ICCV 2023 1 · Xiaotian Li, Xiang Zhang, Taoyue Wang, Lijun Yin

Recent studies on dynamic facial action unit (AU) detection have extensively relied on dense annotations. However, manual annotations are difficult, time-consuming, and costly. The canonical semi-supervised learning …

Knowledge Distillation

Leveraging Action Affinity and Continuity for Semi-supervised Temporal Action Segmentation

2022-07-18 · Guodong Ding, Angela Yao

We present a semi-supervised learning approach to the temporal action segmentation task. The goal of the task is to temporally detect and segment actions in long, untrimmed procedural videos, where only a small set of vi…

Action SegmentationTemporal Action Segmentation