paper-with-me

홈 › Papers

UTAL-GNN: Unsupervised Temporal Action Localization using Graph Neural Networks

2025-08-27 · Bikash Kumar Badatya, Vipul Baghel, Ravi Hegde arxiv

Fine-grained action localization in untrimmed sports videos presents a significant challenge due to rapid and subtle motion transitions over short durations. Existing supervised and weakly supervised solutions often rely on extensive annotated datasets and high-capacity models, making them computationally intensive and less adaptable to real-world scenarios. In this work, we introduce a lightweight and unsupervised skeleton-based action localization pipeline that leverages spatio-temporal graph neural representations. Our approach pre-trains an Attention-based Spatio-Temporal Graph Convolutional Network (ASTGCN) on a pose-sequence denoising task with blockwise partitions, enabling it to learn intrinsic motion dynamics without any manual labeling. At inference, we define a novel Action Dynamics Metric (ADM), computed directly from low-dimensional ASTGCN embeddings, which detects motion boundaries by identifying inflection points in its curvature profile. Our method achieves a mean Average Precision (mAP) of 82.66% and average localization latency of 29.09 ms on the DSV Diving dataset, matching state-of-the-art supervised performance while maintaining computational efficiency. Furthermore, it generalizes robustly to unseen, in-the-wild diving footage without retraining, demonstrating its practical applicability for lightweight, real-time action analysis systems in embedded or dynamic environments.

📄 PDF Abstract BibTeX arXiv:2508.19647

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal Action LocalizationComputational Efficiency

Similar Papers 제목 키워드 기반

CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization

2025-05-29 · Rui Xia, Dan Jiang, Quan Zhang, Ke Zhang 외

Temporal Action Localization (TAL) has garnered significant attention in information retrieval. Existing supervised or weakly supervised methods heavily rely on labeled temporal boundaries and action categories, which ar…

Action LocalizationInformation RetrievalTemporal Action Localization

BLAB: Brutally Long Audio Bench

2025-05-05 · Orevaoghene Ahia, Martijn Bartelds, Kabir Ahuja, Hila Gonen 외

Developing large audio language models (LMs) capable of understanding diverse spoken interactions is essential for accommodating the multimodal nature of human communication and can increase the accessibility of language…

Form

Clip-level Uncertainty and Temporal-aware Active Learning for End-to-End Multi-Object Tracking

2026-05-11 · Riku Inoue, Shogo Sato, Kazuhiko Murasaki, Tomoyasu Shimada 외 arxiv

Multi-Object Tracking (MOT) in dynamic environments relies on robust temporal reasoning to maintain consistent object identities over time. Transformer-based end-to-end MOT models achieve strong performance by explicitly…

Multi-Object TrackingActive Learning

Unsupervised Pre-training for Temporal Action Localization Tasks

2022-03-25 · CVPR 2022 1 · Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao 외

Unsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pre-trained models can be sub-optimal…

Action LocalizationContrastive LearningRepresentation LearningTemporal Action Localization+3

Learning Temporal Co-Attention Models for Unsupervised Video Action Localization

2020-06-01 · CVPR 2020 6 · Guoqiang Gong, Xinghan Wang, Yadong Mu, Qi Tian

Temporal action localization (TAL) in untrimmed videos recently receives tremendous research enthusiasm. To our best knowledge, this is the first attempt in the literature to explore this task under an unsupervised setti…

Action LocalizationClusteringTemporal Action LocalizationTriplet