paper-with-me

홈 › Papers

The Blessings of Unlabeled Background in Untrimmed Videos

2021-03-24 · CVPR 2021 1 · YuAn Liu, Jingyuan Chen, Zhenfang Chen, Bing Deng, Jianqiang Huang, Hanwang Zhang

Weakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish the action of interest segments from the background, which is unlabelled even on the video-level. While previous works treat the background as "curses", we consider it as "blessings". Specifically, we first use causal analysis to point out that the common localization errors are due to the unobserved confounder that resides ubiquitously in visual recognition. Then, we propose a Temporal Smoothing PCA-based (TS-PCA) deconfounder, which exploits the unlabelled background to model an observed substitute for the unobserved confounder, to remove the confounding effect. Note that the proposed deconfounder is model-agnostic and non-intrusive, and hence can be applied in any WTAL method without model re-designs. Through extensive experiments on four state-of-the-art WTAL methods, we show that the deconfounder can improve all of them on the public datasets: THUMOS-14 and ActivityNet-1.3.

📄 PDF Abstract BibTeX arXiv:2103.13183

Code (1)

liuyuancv/WTAL_blessing 공식 구현

Tasks

Action LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action Localization

Similar Papers 제목 키워드 기반

Exploring Relations in Untrimmed Videos for Self-Supervised Learning

2020-08-06 · Dezhao Luo, Bo Fang, Yu Zhou, Yucan Zhou 외

Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not really self-s…

Action RecognitionChange DetectionRetrievalSelf-Supervised Learning+1

The THUMOS Challenge on Action Recognition for Videos "in the Wild"

2016-04-21 · Haroon Idrees, Amir R. Zamir, Yu-Gang Jiang, Alex Gorban 외

Automatically recognizing and localizing wide ranges of human actions has crucial importance for video understanding. Towards this goal, the THUMOS challenge was introduced in 2013 to serve as a benchmark for action reco…

Action ClassificationAction RecognitionGeneral ClassificationTemporal Action Localization+1

Learning Transferable Self-attentive Representations for Action Recognition in Untrimmed Videos with Weak Supervision

2019-02-20 · Xiao-Yu Zhang, Haichao Shi, Changsheng Li, Kai Zheng 외

Action recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotation…

Action RecognitionAction Recognition In VideosTemporal Action Localization

Learning to Localize Actions from Moments

2020-08-31 · ECCV 2020 8 · Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian 외

With the knowledge of action moments (i.e., trimmed video clips that each contains an action instance), humans could routinely localize an action temporally in an untrimmed video. Nevertheless, most practical methods sti…

Action LocalizationTransfer Learning

Enabling Weakly-Supervised Temporal Action Localization from On-Device Learning of the Video Stream

2022-08-25 · Yue Tang, Yawen Wu, Peipei Zhou, Jingtong Hu

Detecting actions in videos have been widely applied in on-device applications. Practical on-device videos are always untrimmed with both action and background. It is desirable for a model to both recognize the class of …

Action DetectionAction LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action Localization