paper-with-me

홈 › Papers

Multidomain Multimodal Fusion For Human Action Recognition Using Inertial Sensors

2020-08-22 · Zeeshan Ahmad, Naimul Khan

One of the major reasons for misclassification of multiplex actions during action recognition is the unavailability of complementary features that provide the semantic information about the actions. In different domains these features are present with different scales and intensities. In existing literature, features are extracted independently in different domains, but the benefits from fusing these multidomain features are not realized. To address this challenge and to extract complete set of complementary information, in this paper, we propose a novel multidomain multimodal fusion framework that extracts complementary and distinct features from different domains of the input modality. We transform input inertial data into signal images, and then make the input modality multidomain and multimodal by transforming spatial domain information into frequency and time-spectrum domain using Discrete Fourier Transform (DFT) and Gabor wavelet transform (GWT) respectively. Features in different domains are extracted by Convolutional Neural networks (CNNs) and then fused by Canonical Correlation based Fusion (CCF) for improving the accuracy of human action recognition. Experimental results on three inertial datasets show the superiority of the proposed method when compared to the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2008.09748

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

2025-07-17 · Penglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 외

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, …

Question AnsweringScene Understanding

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

2025-12-04 · Novanto Yudistira arxiv

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and…

Self-Supervised LearningAction RecognitionAction Detection

Multimodal Language Analysis with Recurrent Multistage Fusion

2018-08-12 · EMNLP 2018 10 · Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities. Comprehending multimodal language requires modeling n…

Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition

2025-12-26 · Zeyu Liang, Hailun Xia, Naichuan Zheng arxiv

While human action recognition has witnessed notable achievements, multimodal methods fusing RGB and skeleton modalities still suffer from their inherent heterogeneity and fail to fully exploit the complementary potentia…

Graph Representation LearningAction Recognition

Vision and Inertial Sensing Fusion for Human Action Recognition : A Review

2020-08-02 · Sharmin Majumder, Nasser Kehtarnavaz

Human action recognition is used in many applications such as video surveillance, human computer interaction, assistive living, and gaming. Many papers have appeared in the literature showing that the fusion of vision an…

Action RecognitionTemporal Action Localization