paper-with-me

홈 › Papers

Video2IMU: Realistic IMU features and signals from videos

2022-02-14 · Arttu Lämsä, Jaakko Tervonen, Jussi Liikka, Constantino Álvarez Casado, Miguel Bordallo López

Human Activity Recognition (HAR) from wearable sensor data identifies movements or activities in unconstrained environments. HAR is a challenging problem as it presents great variability across subjects. Obtaining large amounts of labelled data is not straightforward, since wearable sensor signals are not easy to label upon simple human inspection. In our work, we propose the use of neural networks for the generation of realistic signals and features using human activity monocular videos. We show how these generated features and signals can be utilized, instead of their real counterparts, to train HAR models that can recognize activities using signals obtained with wearable sensors. To prove the validity of our methods, we perform experiments on an activity recognition dataset created for the improvement of industrial work safety. We show that our model is able to realistically generate virtual sensor signals and features usable to train a HAR classifier with comparable performance as the one trained using real sensor data. Our results enable the use of available, labelled video data for training HAR models to classify signals from wearable sensors.

📄 PDF Abstract BibTeX arXiv:2202.06547

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionHuman Activity Recognition

Similar Papers 제목 키워드 기반

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

2023-03-29 · Jiawei Liu, Weining Wang, Sihan Chen, Xinxin Zhu 외

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic vide…

Audio GenerationContrastive LearningDecoderVideo Generation

Realistic Speech-Driven Facial Animation with GANs

2019-06-14 · Konstantinos Vougioukas, Stavros Petridis, Maja Pantic

Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This …

Audio-Visual SynchronizationLip Reading

End-to-End Speech-Driven Facial Animation with Temporal GANs

2018-05-23 · Konstantinos Vougioukas, Stavros Petridis, Maja Pantic

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This …

Lip Reading

Inserting Videos into Videos

2019-03-15 · CVPR 2019 6 · Donghoon Lee, Tomas Pfister, Ming-Hsuan Yang

In this paper, we introduce a new problem of manipulating a given video by inserting other videos into it. Our main task is, given an object video and a scene video, to insert the object video at a user-specified locatio…

ObjectObject TrackingPerson Re-Identification

Editing Physiological Signals in Videos Using Latent Representations

2025-09-29 · Tianwen Zhou, Akshay Paruchuri, Josef Spjut, Kaan Akşit arxiv

Camera-based physiological signal estimation provides a non-contact and convenient means to monitor Heart Rate (HR). However, the presence of vital signals in facial videos raises significant privacy concerns, as they ca…