paper-with-me

Papers

Unsupervised Discriminative Embedding for Sub-Action Learning in Complex Activities

2021-04-30 · Sirnam Swetha, Hilde Kuehne, Yogesh S Rawat, Mubarak Shah

Action recognition and detection in the context of long untrimmed video sequences has seen an increased attention from the research community. However, annotation of complex activities is usually time consuming and challenging in practice. Therefore, recent works started to tackle the problem of unsupervised learning of sub-actions in complex activities. This paper proposes a novel approach for unsupervised sub-action learning in complex activities. The proposed method maps both visual and temporal representations to a latent space where the sub-actions are learnt discriminatively in an end-to-end fashion. To this end, we propose to learn sub-actions as latent concepts and a novel discriminative latent concept learning (DLCL) module aids in learning sub-actions. The proposed DLCL module lends on the idea of latent concepts to learn compact representations in the latent embedding space in an unsupervised way. The result is a set of latent vectors that can be interpreted as cluster centers in the embedding space. The latent space itself is formed by a joint visual and temporal embedding capturing the visual similarity and temporal ordering of the data. Our joint learning with discriminative latent concept module is novel which eliminates the need for explicit clustering. We validate our approach on three benchmark datasets and show that the proposed combination of visual-temporal embedding and discriminative latent concepts allow to learn robust action representations in an unsupervised setting.

📄 PDF Abstract BibTeX arXiv:2105.00067

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAction Segmentation

Similar Papers 제목 키워드 기반

Unsupervised Learning and Segmentation of Complex Activities from Video

2018-03-26 · CVPR 2018 6 · Fadime Sener, Angela Yao

This paper presents a new method for unsupervised segmentation of complex activities from video into multiple steps, or sub-activities, without any textual input. We propose an iterative discriminative-generative approac…

Discriminative Hierarchical Modeling of Spatio-Temporally Composable Human Activities

2014-06-01 · CVPR 2014 6 · Ivan Lillo, Alvaro Soto, Juan Carlos Niebles

This paper proposes a framework for recognizing complex human activities in videos. Our method describes human activities in a hierarchical discriminative model that operates at three semantic levels. At the lower level,…

Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed Sequences

2020-01-29 · Rosaura G. VidalMata, Walter J. Scheirer, Anna Kukleva, David Cox 외

Understanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- …

Action RecognitionAction SegmentationTemporal Localization

Unsupervised Embedding Learning for Human Activity Recognition Using Wearable Sensor Data

2023-07-21 · Taoran Sheng, Manfred Huber

The embedded sensors in widely used smartphones and other wearable devices make the data of human activities more accessible. However, recognizing different human activities from the wearable sensor data remains a challe…

Activity RecognitionClusteringHuman Activity Recognition

The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities

2014-06-01 · CVPR 2014 6 · Hilde Kuehne, Ali Arslan, Thomas Serre

This paper describes a framework for modeling human activities as temporally structured processes. Our approach is motivated by the inherently hierarchical nature of human activities and the close correspondence between …

Semantic Parsingspeech-recognitionSpeech Recognition