Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition
Sensor data streams provide valuable information around activities and context for downstream applications, though integrating complementary information can be challenging. We show that large language models (LLMs) can be used for late fusion for activity classification from audio and motion time series data. We curated a subset of data for diverse activity recognition across contexts (e.g., household activities, sports) from the Ego4D dataset. Evaluated LLMs achieved 12-class zero- and one-shot classification F1-scores significantly above chance, with no task-specific training. Zero-shot classification via LLM-based fusion from modality-specific models can enable multimodal temporal applications where there is limited aligned training data for learning a shared embedding space. Additionally, LLM-based fusion can enable model deploying without requiring additional memory and computation for targeted application-specific multimodal models.
Code (0)
등록된 구현이 없습니다.
Tasks
Activity RecognitionSimilar Papers 제목 키워드 기반
MuMu: Cooperative Multitask Learning-based Guided Multimodal Fusion
Multimodal sensors (visual, non-visual, and wearable) can provide complementary information to develop robust perception systems for recognizing activities accurately. However, it is challenging to extract robust multimo…
Activity RecognitionHuman Activity RecognitionMultimodal Activity RecognitionRobust Multimodal Fusion for Human Activity Recognition
The proliferation of IoT and mobile devices equipped with heterogeneous sensors has enabled new applications that rely on the fusion of time-series data generated by multiple sensors with different modalities. While ther…
Activity RecognitionDenoisingHuman Activity RecognitionTime Series+1EmbraceNet for Activity: A Deep Multimodal Fusion Architecture for Activity Recognition
Human activity recognition using multiple sensors is a challenging but promising task in recent decades. In this paper, we propose a deep multimodal fusion model for activity recognition based on the recently proposed fe…
Activity RecognitionHuman Activity RecognitionEgocentric Activity Recognition with Multimodal Fisher Vector
With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egoc…
Activity RecognitionEgocentric Activity RecognitionSelf-Supervised Multimodal Fusion Transformer for Passive Activity Recognition
The pervasiveness of Wi-Fi signals provides significant opportunities for human sensing and activity recognition in fields such as healthcare. The sensors most commonly used for passive Wi-Fi sensing are based on passive…
Activity RecognitionSelf-Supervised LearningSensor Fusion