paper-with-me

Papers

Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations

2024-02-10 · Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

To understand why self-supervised learning (SSL) models have empirically achieved strong performances on several speech-processing downstream tasks, numerous studies have focused on analyzing the encoded information of the SSL layer representations in adult speech. Limited work has investigated how pre-training and fine-tuning affect SSL models encoding children's speech and vocalizations. In this study, we aim to bridge this gap by probing SSL models on two relevant downstream tasks: (1) phoneme recognition (PR) on the speech of adults, older children (8-10 years old), and younger children (1-4 years old), and (2) vocalization classification (VC) distinguishing cry, fuss, and babble for infants under 14 months old. For younger children's PR, the superiority of fine-tuned SSL models is largely due to their ability to learn features that represent older children's speech and then adapt those features to the speech of younger children. For infant VC, SSL models pre-trained on large-scale home recordings learn to leverage phonetic representations at middle layers, and thereby enhance the performance of this task.

📄 PDF Abstract BibTeX arXiv:2402.06888

Code (0)

등록된 구현이 없습니다.

Tasks

Phoneme RecognitionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Improving Children's Speech Recognition by Fine-tuning Self-supervised Adult Speech Representations

2022-11-14 · Renee Lu, Mostafa Shahin, Beena Ahmed

Children's speech recognition is a vital, yet largely overlooked domain when building inclusive speech technologies. The major challenge impeding progress in this domain is the lack of adequate child speech corpora; howe…

Self-Supervised Learningspeech-recognitionSpeech Recognition

Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors

2025-02-12 · Vishwanath Pratap Singh, Md. Sahidullah, Tomi Kinnunen

The increasing use of children's automatic speech recognition (ASR) systems has spurred research efforts to improve the accuracy of models designed for children's speech in recent years. The current approach utilizes eit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingCausal Inference+3

Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech

2025-08-14 · Abhijit Sinha, Harishankar Kumar, Mohit Joshi, Hemant Kumar Kathania 외 arxiv

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech t…

Age And Gender ClassificationSelf-Supervised Learning

Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?

2025-08-28 · Abhijit Sinha, Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Shrikanth Narayanan arxiv

Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervi…

Self-Supervised LearningSpeech Recognition

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

2026-06-20 · Abhijit Sinha, Hemant Kumar Kathania, Mohit Joshi, Harishankar Kumar 외 arxiv

Self-supervised learning (SSL) models have become a central component of modern speech processing systems, as they enable the learning of rich acoustic representations without reliance on labeled data. Despite their succ…

Age And Gender ClassificationSelf-Supervised Learning