paper-with-me

Papers

MuMu: Cooperative Multitask Learning-based Guided Multimodal Fusion

2022-02-22 · AAAI 2022 2 · Md Mofijul Islam, Tariq Iqbal

Multimodal sensors (visual, non-visual, and wearable) can provide complementary information to develop robust perception systems for recognizing activities accurately. However, it is challenging to extract robust multimodal representations due to the heterogeneous characteristics of data from multimodal sensors and disparate human activities, especially in the presence of noisy and misaligned sensor data. In this work, we propose a cooperative multitask learning-based guided multimodal fusion approach, MuMu, to extract robust multimodal representations for human activity recognition (HAR). MuMu employs an auxiliary task learning approach to extract features specific to each set of activities with shared characteristics (activity-group). MuMu then utilizes activity-group-specific features to direct our proposed Guided Multimodal Fusion Approach (GM-Fusion) for extracting complementary multimodal representations, designed as the target task. We evaluated MuMu by comparing its performance to state-of-the-art multimodal HAR approaches on three activity datasets. Our extensive experimental results suggest that MuMu outperforms all the evaluated approaches across all three datasets. Additionally, the ablation study suggests that MuMu significantly outperforms the baseline models (p<0.05), which do not use our guided multimodal fusion. Finally, the robust performance of MuMu on noisy and misaligned sensor data posits that our approach is suitable for HAR in real-world settings.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionHuman Activity RecognitionMultimodal Activity Recognition

Similar Papers 제목 키워드 기반

MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data

2024-06-26 · William Berman, Alexander Peysakhovich

We train a model to generate images from multimodal prompts of interleaved text and images such as "a <picture of a man> man and his <picture of a dog> dog in an <picture of a cartoon> animated style." We bootstrap a mul…

DecoderGPUImage CaptioningImage Generation+3

Channel Exchanging Networks for Multimodal and Multitask Dense Image Prediction

2021-12-04 · Yikai Wang, Fuchun Sun, Wenbing Huang, Fengxiang He 외

Multimodal fusion and multitask learning are two vital topics in machine learning. Despite the fruitful progress, existing methods for both problems are still brittle to the same challenge -- it remains dilemmatic to int…

Semantic Segmentation

MuMUR : Multilingual Multimodal Universal Retrieval

2022-08-24 · Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan, Shachar Rosenman 외

Multi-modal retrieval has seen tremendous progress with the development of vision-language models. However, further improving these models require additional labelled data which is a huge manual effort. In this paper, we…

Image RetrievalMachine TranslationRetrievalTransfer Learning+1

Decentralized Multitask Learning over Learned Task Graphs

2026-08-27 · Zirui Wan, Stefan Vlaski arxiv

This paper investigates decentralized multitask learning over networks when the underlying task relationships are unknown. While existing graph-regularized multitask frameworks typically assume a known structure, practic…

FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction

2024-06-17 · Muhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng 외

Multimodal electronic health record (EHR) data can offer a holistic assessment of a patient's health status, supporting various predictive healthcare tasks. Recently, several studies have embraced the multitask learning …