paper-with-me

Papers

Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech

2025-08-14 · Abhijit Sinha, Harishankar Kumar, Mohit Joshi, Hemant Kumar Kathania, Shrikanth Narayanan, Sudarsana Reddy Kadiri arxiv

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.

📄 PDF Abstract BibTeX arXiv:2508.10332

Code (0)

등록된 구현이 없습니다.

Tasks

Age And Gender ClassificationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition

2024-02-04 · Alexandra Saliba, Yuanchao Li, Ramon Sanabria, Catherine Lai

The efficacy of self-supervised speech models has been validated, yet the optimal utilization of their representations remains challenging across diverse tasks. In this study, we delve into Acoustic Word Embeddings (AWEs…

Emotion RecognitionSpeech Emotion RecognitionWord Embeddings

Ensemble knowledge distillation of self-supervised speech models

2023-02-24 · Kuan-Po Huang, Tzu-hsun Feng, Yu-Kuan Fu, Tsu-Yuan Hsu 외

Distilled self-supervised models have shown competitive performance and efficiency in recent years. However, there is a lack of experience in jointly distilling multiple self-supervised speech models. In our work, we per…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionKnowledge Distillation+4

Comparative layer-wise analysis of self-supervised speech models

2022-11-08 · Ankita Pasad, Bowen Shi, Karen Livescu

Many self-supervised speech models, varying in their pre-training objective, input modality, and pre-training data, have been proposed in the last few years. Despite impressive successes on downstream tasks, we still hav…

speech-recognitionSpeech RecognitionSpoken Language Understanding

Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

2026-05-04 · Sandra Arcos-Holzinger, Sarah M. Erfani, James Bailey, Sanjeev Khudanpur arxiv

Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation sim…

Speech RecognitionAnomaly Detection

A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models

2024-08-24 · Antón de la Fuente, Dan Jurafsky

This study asks how self-supervised speech models represent suprasegmental categories like Mandarin lexical tone, English lexical stress, and English phrasal accents. Through a series of probing tasks, we make layer-wise…

Specificity