paper-with-me

Papers

Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis

2023-09-13 · Jialu Li, Mark Hasegawa-Johnson, Karrie Karahalios

The assessment of children at risk of autism typically involves a clinician observing, taking notes, and rating children's behaviors. A machine learning model that can label adult and child audio may largely save labor in coding children's behaviors, helping clinicians capture critical events and better communicate with parents. In this study, we leverage Wav2Vec 2.0 (W2V2), pre-trained on 4300-hour of home audio of children under 5 years old, to build a unified system for tasks of clinician-child speaker diarization and vocalization classification (VC). To enhance children's VC, we build a W2V2 phoneme recognition system for children under 4 years old, and we incorporate its phonetically-tuned embeddings as auxiliary features or recognize pseudo phonetic transcripts as an auxiliary task. We test our method on two corpora (Rapid-ABC and BabbleCor) and obtain consistent improvements. Additionally, we outperform the state-of-the-art performance on the reproducible subset of BabbleCor. Code available at https://huggingface.co/lijialudew

📄 PDF Abstract BibTeX arXiv:2309.07287

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme RecognitionSelf-Supervised Learningspeaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations

2024-02-10 · Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

To understand why self-supervised learning (SSL) models have empirically achieved strong performances on several speech-processing downstream tasks, numerous studies have focused on analyzing the encoded information of t…

Phoneme RecognitionSelf-Supervised Learning

Employing self-supervised learning models for cross-linguistic child speech maturity classification

2025-06-10 · Theo Zhang, Madurya Suresh, Anne S. Warlaumont, Kasia Hitczenko 외

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art tran…

Self-Supervised Learningvalid

Infrequent Child-Directed Speech Is Bursty and May Draw Infant Vocalizations

2026-03-25 · Margaret Cychosz, Adriana Weisleder arxiv

Children in many parts of the world hear relatively little speech directed to them, yet still reach major language development milestones. What differs about the speech input that infants learn from when directed input i…

Linguistic Input and Child Vocalization of 7 Children from 5 to 30 Months: A Longitudinal Study with LENA Automatic Analysis

2020-06-01 · IJCLCLP 2020 6 · Chia-Cheng Lee, Li-mei Chen, D. Kimbrough Oller

Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations

2024-07-08 · Guillem Bonafos, Clara Bourot, Pierre Pudlo, Jean-Marc Freyermuth 외

Based on audio recordings made once a month during the first 12 months of a child's life, we propose a new method for clustering this set of vocalizations. We use a topologically augmented representation of the vocalizat…