Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multimodal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionActivity RecognitionDomain GeneralizationSimilar Papers 제목 키워드 기반
EgoAVU: Egocentric Audio-Visual Understanding
Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labe…
QuerYD: A video dataset with high-quality text and audio narrations
We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality…
RetrievalVideo UnderstandingVocal Bursts Intensity PredictionNymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data
The Nymeria Dataset, released in 2024, is a large-scale collection of in-the-wild human activities captured with multiple egocentric wearable devices that are spatially localized and temporally synchronized. It provides …
Point CloudsLearning to Generate Long-term Future Narrations Describing Activities of Daily Living
Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's futur…
Action AnticipationDecision MakingLanguage ModelingLanguage Modelling+1Prosody Analysis of Audiobooks
Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…
AttributeLanguage ModelingLanguage ModellingProsody Prediction+2