Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
Multimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variational Auto-Encoder (VQ-VAE) based Multimodal Representation Learning (MRL) model to produce task-agnostic speech representations from vocal Tract Variables (TVs) and Facial Action Units (FAUs). These representations are then used in a Multi-Task Learning (MTL) based downstream prediction model to obtain class labels and an overall severity score. The proposed framework outperforms the previous works on the multi-class classification task across all evaluation metrics (Weighted F1 score, AUC-ROC score, and Weighted Accuracy). Additionally, it estimates the schizophrenia severity score, a task not addressed by earlier approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-class ClassificationMulti-Task LearningRepresentation LearningSimilar Papers 제목 키워드 기반
A study on the impact of Self-Supervised Learning on automatic dysarthric speech assessment
Automating dysarthria assessments offers the opportunity to develop practical, low-cost tools that address the current limitations of manual and subjective assessments. Nonetheless, the small size of most dysarthria data…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationSelf-Supervised Learning+2Automatic Pronunciation Assessment using Self-Supervised Speech Representation Learning
Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community. In particular, speech representations learned by SSL model…
Representation LearningSelf-Supervised LearningSpeech Representation LearningNon Intrusive Intelligibility Predictor for Hearing Impaired Individuals using Self Supervised Speech Representations
Self-supervised speech representations (SSSRs) have been successfully applied to a number of speech-processing tasks, e.g. as feature extractor for speech quality (SQ) prediction, which is, in turn, relevant for assessme…
PredictionSpeech EnhancementL2 proficiency assessment using self-supervised speech representations
There has been a growing demand for automated spoken language assessment systems in recent years. A standard pipeline for this process is to start with a speech recognition system and derive features, either hand-crafted…
speech-recognitionSpeech RecognitionLearning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…
Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1