paper-with-me

홈 › Papers

RNNs, CNNs and Transformers in Human Action Recognition: A Survey and a Hybrid Model

2024-06-02 · Khaled Alomar, Halil Ibrahim Aysel, Xiaohao Cai

Human Action Recognition (HAR) encompasses the task of monitoring human activities across various domains, including but not limited to medical, educational, entertainment, visual surveillance, video retrieval, and the identification of anomalous activities. Over the past decade, the field of HAR has witnessed substantial progress by leveraging Convolutional Neural Networks (CNNs) to effectively extract and comprehend intricate information, thereby enhancing the overall performance of HAR systems. Recently, the domain of computer vision has witnessed the emergence of Vision Transformers (ViTs) as a potent solution. The efficacy of transformer architecture has been validated beyond the confines of image analysis, extending their applicability to diverse video-related tasks. Notably, within this landscape, the research community has shown keen interest in HAR, acknowledging its manifold utility and widespread adoption across various domains. This article aims to present an encompassing survey that focuses on CNNs and the evolution of Recurrent Neural Networks (RNNs) to ViTs given their importance in the domain of HAR. By conducting a thorough examination of existing literature and exploring emerging trends, this study undertakes a critical analysis and synthesis of the accumulated knowledge in this field. Additionally, it investigates the ongoing efforts to develop hybrid approaches. Following this direction, this article presents a novel hybrid model that seeks to integrate the inherent strengths of CNNs and ViTs.

📄 PDF Abstract BibTeX arXiv:2407.06162

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action LocalizationVideo Retrieval

Similar Papers 제목 키워드 기반

Deep Learning Approaches for Human Action Recognition in Video Data

2024-03-11 · Yufei Xie

Human action recognition in videos is a critical task with significant implications for numerous applications, including surveillance, sports analytics, and healthcare. The challenge lies in creating models that are both…

Action RecognitionAction Recognition In VideosDeep LearningSports Analytics+2

Lattice Long Short-Term Memory for Human Action Recognition

2017-08-13 · ICCV 2017 10 · Lin Sun, Kui Jia, Kevin Chen, Dit Yan Yeung 외

Human actions captured in video sequences are three-dimensional signals characterizing visual appearance and motion dynamics. To learn action patterns, existing methods adopt Convolutional and/or Recurrent Neural Network…

Action RecognitionOptical Flow EstimationTemporal Action Localization

Making Convolutional Networks Recurrent for Visual Sequence Learning

2018-06-01 · CVPR 2018 6 · Xiaodong Yang, Pavlo Molchanov, Jan Kautz

Recurrent neural networks (RNNs) have emerged as a powerful model for a broad range of machine learning problems that involve sequential data. While an abundance of work exists to understand and improve RNNs in the conte…

Action RecognitionFace AlignmentGesture RecognitionHand Gesture Recognition+6

Lightweight Transformer in Federated Setting for Human Activity Recognition

2021-10-01 · Ali Raza, Kim Phuc Tran, Ludovic Koehl, Shujun Li 외

Human activity recognition (HAR) is a machine learning task with important applications in healthcare especially in the context of home care of patients and older adults. HAR is often based on data collected from smart s…

Activity RecognitionFederated LearningHuman Activity Recognition

Memory-Augmented Temporal Dynamic Learning for Action Recognition

2019-04-30 · Yuan Yuan, Dong Wang, Qi. Wang

Human actions captured in video sequences contain two crucial factors for action recognition, i.e., visual appearance and motion dynamics. To model these two aspects, Convolutional and Recurrent Neural Networks (CNNs and…

Action RecognitionTemporal Action Localization