paper-with-me

홈 › Papers

Action-Agnostic Point-Level Supervision for Temporal Action Detection

2024-12-30 · Shuhei M. Yoshida, Takashi Shibata, Makoto Terao, Takayuki Okatani, Masashi Sugiyama

We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS '14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.

📄 PDF Abstract BibTeX arXiv:2412.21205

Code (1)

smy-nec/aapl 공식 구현

Tasks

Action Detection

Similar Papers 제목 키워드 기반

Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization

2026-02-05 · Yunchuan Ma, Laiyun Qing, Guorong Li, Yuqing Liu 외 arxiv

Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per action instance) to train a model to effectively locate action instances wit…

Weakly-supervised Temporal Action LocalizationMulti-Task Learning

Point-Level Temporal Action Localization: Bridging Fully-supervised Proposals to Weakly-supervised Losses

2020-12-15 · Chen Ju, Peisen Zhao, Ya zhang, Yanfeng Wang 외

Point-Level temporal action localization (PTAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the frame-level prediction paradigm to lear…

Action LocalizationPredictionTemporal Action LocalizationWeakly Supervised Action Localization

Weakly Supervised Action Localization by Sparse Temporal Pooling Network

2017-12-14 · CVPR 2018 6 · Phuc Nguyen, Ting Liu, Gautam Prasad, Bohyung Han

We propose a weakly supervised temporal action localization algorithm on untrimmed videos using convolutional neural networks. Our algorithm learns from video-level class labels and predicts temporal intervals of human a…

Action ClassificationAction LocalizationTemporal Action LocalizationTemporal Localization+2

Proposal-based Temporal Action Localization with Point-level Supervision

2023-10-09 · Yuan Yin, Yifei HUANG, Ryosuke Furuta, Yoichi Sato

Point-level supervised temporal action localization (PTAL) aims at recognizing and localizing actions in untrimmed videos where only a single point (frame) within every action instance is annotated in training data. With…

Action ClassificationAction LocalizationMultiple Instance LearningTemporal Action Localization

PointAction: 3D Points as Universal Action Representations for Robot Control

2026-06-02 · Mutian Tong, Han Jiang, Qiao Feng, Lingjie Liu 외 arxiv

Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not di…

Robot ManipulationVideo GenerationVideo Prediction