paper-with-me

홈 › Papers

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

2026-06-26 · Ahmed Mohamady, Robin Burchard, Kristof Van Laerhoven arxiv

Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their uni-modal counterparts. Modalities can include IMUs, RGB cameras, audio signals, and others. One important aspect of multi-modal deep learning is the sensor fusion approach we apply. Over recent years, multiple fusion paradigms have been proposed for multi-modal HAR. However, to the best of our knowledge, no head-to-head comparison of these paradigms exists on a common multi-modal HAR benchmark dataset. To address this research gap, we systematically compare seven state-of-the-art sensor fusion methods on the recently released HARMES dataset, which comprises 61 hours of fully labeled IMU, audio, and ambient humidity data. The chosen dataset focuses on 15 household and personal hygiene activities of daily living (ADLs). By applying the seven different fusion techniques to a state-of-the-art multi-modal model architecture, we show that Gated Multi-modal Fusion achieves the highest macro F1-score (0.82), surpassing the concatenation-based late fusion HARMES paper baseline of 0.76 by +6pp under leave-one-participant-out evaluation. All code used in our experiments is made publicly available on GitHub.

📄 PDF Abstract BibTeX arXiv:2606.27886

Code (0)

등록된 구현이 없습니다.

Tasks

Human Activity Recognition

Similar Papers 제목 키워드 기반

Multimodal Object Detection in Remote Sensing

2023-07-13 · Abdelbadie Belmouhcine, Jean-Christophe Burnel, Luc Courtrai, Minh-Tan Pham 외

Object detection in remote sensing is a crucial computer vision task that has seen significant advancements with deep learning techniques. However, most existing works in this area focus on the use of generic object dete…

Objectobject-detectionObject DetectionSurvey

Application of Multimodal Fusion Deep Learning Model in Disease Recognition

2024-05-22 · Xiaoyi Liu, Hongjie Qiu, Muqing Li, Zhou Yu 외

This paper introduces an innovative multi-modal fusion deep learning approach to overcome the drawbacks of traditional single-modal recognition techniques. These drawbacks include incomplete information and limited diagn…

Deep LearningDiagnostic

Multi-Modal Deep Learning for Credit Rating Prediction Using Text and Numerical Data Streams

2023-04-21 · Mahsa Tavakoli, Rohitash Chandra, Fengrui Tian, Cristián Bravo

Knowing which factors are significant in credit rating assignment leads to better decision-making. However, the focus of the literature thus far has been mostly on structured data, and fewer studies have addressed unstru…

Decision MakingDeep Learning

Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

2018-07-01 · ACL 2018 7 · AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria 외

Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically this language is multimodal (heterogeneous), sequential and asynchronous; it consists of the language (words), visual (expressions…

Emotion RecognitionLanguage ModelingLanguage ModellingMultimodal Sentiment Analysis+2

Hybrid Attention based Multimodal Network for Spoken Language Classification

2018-08-01 · COLING 2018 8 · Yue Gu, Kangning Yang, Shiyu Fu, Shuhong Chen 외

We examine the utility of linguistic content and vocal characteristics for multimodal deep learning in human spoken language understanding. We present a deep multimodal network with both feature attention and modality at…

ClassificationDeep LearningEmotion RecognitionGeneral Classification+3