paper-with-me

홈 › Papers

Self-supervised representations in speech-based depression detection

2023-05-20 · Wen Wu, Chao Zhang, Philip C. Woodland

This paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived from different layers of pre-trained foundation models is first presented for SDD, which provides insight to suitable indicator for depression detection. Knowledge transfer is then performed from automatic speech recognition (ASR) and emotion recognition to SDD by fine-tuning the foundation models. Results show that the uses of oracle and ASR transcriptions yield similar SDD performance when the hidden representations of the ASR model is incorporated along with the ASR textual information. By integrating representations from multiple foundation models, state-of-the-art SDD results based on real ASR were achieved on the DAIC-WOZ dataset.

📄 PDF Abstract BibTeX arXiv:2305.12263

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Depression DetectionEmotion RecognitionSelf-Supervised Learningspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

2026-07-03 · Anisha Pattanayak, Huang-Cheng Chou, Shrikanth Narayanan, Sudarsana Reddy Kadiri hf

Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a step called temporal aggregation rarely studied on its own. Most benchmarks fix a single self-supervised …

Why Pre-trained Models Fail: Feature Entanglement in Multi-modal Depression Detection

2025-03-09 · Xiangyu Zhang, Beena Ahmed, Julien Epps

Depression remains a pressing global mental health issue, driving considerable research into AI-driven detection approaches. While pre-trained models, particularly speech self-supervised models (SSL Models), have been ap…

Data AugmentationDepression Detection

Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech

2025-10-05 · Yuxin Li, Eng Siong Chng, Cuntai Guan arxiv

Speech-based depression detection (SDD) has emerged as a non-invasive and scalable alternative to conventional clinical assessments. However, existing methods still struggle to capture robust depression-related speech ch…

Self-Supervised LearningRepresentation Learning

Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection

2026-03-03 · Kashaf Gulzar, Korbinian Riedhammer, Elmar Nöth, Andreas K. Maier 외 arxiv

Speech-based detection of cognitive impairment (CI) offers a promising non-invasive approach for early diagnosis, yet performance disparities across demographic and clinical subgroups remain underexplored, raising concer…

Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection

2023-10-06 · Ziyun Cui, Wen Wu, Wei-Qiang Zhang, Ji Wu 외

The detection of Alzheimer's disease (AD) from spontaneous speech has attracted increasing attention while the sparsity of training data remains an important issue. This paper handles the issue by knowledge transfer, spe…

Alzheimer's Disease DetectionDepression DetectionTransfer Learning