Non-uniform Speaker Disentanglement For Depression Detection From Raw Speech Signals
While speech-based depression detection methods that use speaker-identity features, such as speaker embeddings, are popular, they often compromise patient privacy. To address this issue, we propose a speaker disentanglement method that utilizes a non-uniform mechanism of adversarial SID loss maximization. This is achieved by varying the adversarial weight between different layers of a model during training. We find that a greater adversarial weight for the initial layers leads to performance improvement. Our approach using the ECAPA-TDNN model achieves an F1-score of 0.7349 (a 3.7% improvement over audio-only SOTA) on the DAIC-WoZ dataset, while simultaneously reducing the speaker-identification accuracy by 50%. Our findings suggest that identifying depression through speech signals can be accomplished without placing undue reliance on a speaker's identity, paving the way for privacy-preserving approaches of depression detection.
Code (1)
Tasks
Depression DetectionDisentanglementPrivacy PreservingSpeaker IdentificationSimilar Papers 제목 키워드 기반
A Privacy-Preserving Unsupervised Speaker Disentanglement Method for Depression Detection from Speech
The proposed method focuses on speaker disentanglement in the context of depression detection from speech signals. Previous approaches require patient/speaker labels, encounter instability due to loss maximization, and …
De-identificationDepression DetectionDisentanglementPrivacy PreservingA Step Towards Preserving Speakers' Identity While Detecting Depression Via Speaker Disentanglement
Preserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics…
Depression DetectionDisentanglementDepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection
Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encour…
Significance of Speaker Embeddings and Temporal Context for Depression Detection
Depression detection from speech has attracted a lot of attention in recent years. However, the significance of speaker-specific information in depression detection has not yet been explored. In this work, we analyze the…
Depression DetectionLanguage-Agnostic Analysis of Speech Depression Detection
The people with Major Depressive Disorder (MDD) exhibit the symptoms of tonal variations in their speech compared to the healthy counterparts. However, these tonal variations not only confine to the state of MDD but also…
Depression Detection