paper-with-me

홈 › Papers

Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition

2024-09-15 · Cagri Gungor, Adriana Kovashka

First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multimodal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations.

📄 PDF Abstract BibTeX arXiv:2409.09611

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionActivity RecognitionDomain Generalization

Similar Papers 제목 키워드 기반

EgoAVU: Egocentric Audio-Visual Understanding

2026-02-05 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labe…

QuerYD: A video dataset with high-quality text and audio narrations

2020-11-22 · Andreea-Maria Oncescu, João F. Henriques, Yang Liu, Andrew Zisserman 외

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality…

RetrievalVideo UnderstandingVocal Bursts Intensity Prediction

NymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data

2026-03-19 · Daniel DeTone, Federica Bogo, Eric-Tuan Le, Duncan Frost 외 arxiv

The Nymeria Dataset, released in 2024, is a large-scale collection of in-the-wild human activities captured with multiple egocentric wearable devices that are spatially localized and temporally synchronized. It provides …

Point Clouds

Learning to Generate Long-term Future Narrations Describing Activities of Daily Living

2025-03-03 · Ramanathan Rajendiran, Debaditya Roy, Basura Fernando

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's futur…

Action AnticipationDecision MakingLanguage ModelingLanguage Modelling+1

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2