paper-with-me

홈 › Papers

Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family Audio

2023-05-21 · Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

To perform automatic family audio analysis, past studies have collected recordings using phone, video, or audio-only recording devices like LENA, investigated supervised learning methods, and used or fine-tuned general-purpose embeddings learned from large pretrained models. In this study, we advance the audio component of a new infant wearable multi-modal device called LittleBeats (LB) by learning family audio representation via wav2vec 2.0 (W2V2) pertaining. We show given a limited number of labeled LB home recordings, W2V2 pretrained using 1k-hour of unlabeled home recordings outperforms oracle W2V2 pretrained on 960-hour unlabeled LibriSpeech in terms of parent/infant speaker diarization (SD) and vocalization classifications (VC) at home. Extra relevant external unlabeled and labeled data further benefit W2V2 pretraining and fine-tuning. With SpecAug and environmental speech corruptions, we obtain 12% relative gain on SD and moderate boost on VC. Code and model weights are available.

📄 PDF Abstract BibTeX arXiv:2305.12530

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Visualizations of Complex Sequences of Family-Infant Vocalizations Using Bag-of-Audio-Words Approach Based on Wav2vec 2.0 Features

2022-03-29 · Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

In the U.S., approximately 15-17% of children 2-8 years of age are estimated to have at least one diagnosed mental, behavioral or developmental disorder. However, such disorders often go undiagnosed, and the ability to e…

speaker-diarizationSpeaker Diarization

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

2025-12-11 · Shengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton 외 arxiv

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-lan…

Spatial Reasoning

Measuring Mother-Infant Emotions By Audio Sensing

2019-12-10 · Xuewen Yao, Dong He, Tiancheng Jing, Kaya de Barbaro

It has been suggested in developmental psychology literature that the communication of affect between mothers and their infants correlates with the socioemotional and cognitive development of infants. In this study, we o…

Active Learning

Learning domain-invariant classifiers for infant cry sounds

2023-11-30 · Charles C. Onu, Hemanth K. Sheetha, Arsenii Gorin, Doina Precup

The issue of domain shift remains a problematic phenomenon in most real-world datasets and clinical audio is no exception. In this work, we study the nature of domain shift in a clinical database of infant cry sounds acq…

Domain AdaptationUnsupervised Domain Adaptation

Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry

2026-08-31 · Luigi Simeone arxiv

We compare six self-supervised pretext tasks for infant cry analysis under a fixed budget, meaning the same compact encoder of 1.17M parameters, the same 115 hours of license-verified public pretraining audio, and the sa…

Domain Adaptation