paper-with-me

Papers

Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images

2015-04-04 · Chen Sun, Sanketh Shetty, Rahul Sukthankar, Ram Nevatia

We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos.

📄 PDF Abstract BibTeX arXiv:1504.00983

Code (1)

zhengshou/AutoLoc

Tasks

Action LocalizationAction RecognitionTemporal Action LocalizationTemporal Localization

Similar Papers 제목 키워드 기반

FineAction: A Fine-Grained Video Dataset for Temporal Action Localization

2021-05-24 · Yi Liu, LiMin Wang, Yali Wang, Xiao Ma 외

Temporal action localization (TAL) is an important and challenging problem in video understanding. However, most existing TAL benchmarks are built upon the coarse granularity of action classes, which exhibits two major l…

Action DetectionAction LocalizationFine-Grained Action DetectionTemporal Action Localization+3

Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization

2026-05-22 · Tianyu Wang, Junjie Wu, Jingquan Gao, Shishuo Li arxiv

Temporal Action Localization (TAL) has been extensively studied in generic video understanding, while fine-grained sports scenarios, such as professional badminton, remain underexplored due to their complex and subtle sp…

Temporal Action Localization

Zero-Shot Temporal Action Localization Through Textual Guidance

2026-05-21 · Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Paolo Rota 외 arxiv

Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), …

Temporal Action LocalizationAction Classification

Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos

2022-08-03 · Juncheng Li, Junlin Xie, Linchao Zhu, Long Qian 외

Understanding human emotions is a crucial ability for intelligent robots to provide better human-robot interactions. The existing works are limited to trimmed video-level emotion classification, failing to locate the tem…

Action LocalizationEmotion ClassificationTemporal Action LocalizationWeakly-supervised Learning

Fine-grained Iterative Attention Network for TemporalLanguage Localization in Videos

2020-08-06 · Xiaoye Qu, Pengwei Tang, Zhikang Zhou, Yu Cheng 외

Temporal language localization in videos aims to ground one video segment in an untrimmed video based on a given sentence query. To tackle this task, designing an effective model to extract ground-ing information from bo…

Sentence