Audio-Adaptive Activity Recognition Across Video Domains
This paper strives for activity recognition under domain shift, for example caused by change of scenery or camera viewpoint. The leading approaches reduce the shift in activity appearance by adversarial training and self-supervised learning. Different from these vision-focused works we leverage activity sounds for domain adaptation as they have less variance across domains and can reliably indicate which activities are not happening. We propose an audio-adaptive encoder and associated learning methods that discriminatively adjust the visual feature representation as well as addressing shifts in the semantic distribution. To further eliminate domain-specific features and include domain-invariant activity sounds for recognition, an audio-infused recognizer is proposed, which effectively models the cross-modal interaction across domains. We also introduce the new task of actor shift, with a corresponding audio-visual dataset, to challenge our method with situations where the activity appearance changes dramatically. Experiments on this dataset, EPIC-Kitchens and CharadesEgo show the effectiveness of our approach.
Code (1)
Tasks
Activity RecognitionDomain AdaptationSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Multi-modal Egocentric Activity Recognition using Audio-Visual Features
Egocentric activity recognition in first-person videos has an increasing importance with a variety of applications such as lifelogging, summarization, assisted-living and activity tracking. Existing methods for this task…
Activity RecognitionEgocentric Activity RecognitionOptical Flow EstimationDay2Dark: Pseudo-Supervised Activity Recognition beyond Silent Daylight
This paper strives to recognize activities in the dark, as well as in the day. We first establish that state-of-the-art activity recognizers are effective during the day, but not trustworthy in the dark. The main causes …
Activity RecognitionDomain AdaptationImage EnhancementState of the Art of Audio- and Video-Based Solutions for AAL
The report illustrates the state of the art of the most successful AAL applications and functions based on audio and video data, namely (i) lifelogging and self-monitoring, (ii) remote monitoring of vital signs, (iii) em…
Gesture RecognitionSound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues
Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-perso…
Action RecognitionScene RecognitionVideo AlignmentDynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
Humans do not understand individual events in isolation; rather, they generalize concepts within classes and compare them to others. Existing audio-video pre-training paradigms only focus on the alignment of the overall …
Human Activity Recognition