Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis
Recent works in pathological speech analysis have increasingly relied on powerful self-supervised speech representations, leading to promising results. However, the complex, black-box nature of these embeddings and the limited research on their interpretability significantly restrict their adoption for clinical diagnosis. To address this gap, we propose a novel, interpretable framework specifically designed to support Parkinson's Disease (PD) diagnosis. Through the design of simple yet effective cross-attention mechanisms for both embedding- and temporal-level analysis, the proposed framework offers interpretability from two distinct but complementary perspectives. Experimental findings across five well-established speech benchmarks for PD detection demonstrate the framework's capability to identify meaningful speech patterns within self-supervised representations for a wide range of assessment tasks. Fine-grained temporal analyses further underscore its potential to enhance the interpretability of deep-learning pathological speech models, paving the way for the development of more transparent, trustworthy, and clinically applicable computer-assisted diagnosis systems in this domain. Moreover, in terms of classification accuracy, our method achieves results competitive with state-of-the-art approaches, while also demonstrating robustness in cross-lingual scenarios when applied to spontaneous speech production.
Code (1)
Similar Papers 제목 키워드 기반
ULTra: Unveiling Latent Token Interpretability in Transformer Based Understanding
Transformers have revolutionized Computer Vision (CV) and Natural Language Processing (NLP) through self-attention mechanisms. However, due to their complexity, their latent token representations are often difficult to i…
SegmentationSemantic SegmentationText SummarizationUnsupervised Semantic SegmentationDiscrete Speech Unit Extraction via Independent Component Analysis
Self-supervised speech models (S3Ms) have become a common tool for the speech processing community, leveraging representations for downstream tasks. Clustering S3M representations yields discrete speech units (DSUs), whi…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeech-recognition+1Revisiting transposed convolutions for interpreting raw waveform sound event recognition CNNs by sonification
The majority of recent work on the interpretability of audio and speech processing deep neural networks (DNNs) interprets spectral information modelled by the first layer, relying solely on visual means of interpretation…
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used…
Automatic Speech RecognitionData AugmentationGPUPhoneme Recognition+3Evaluating context-invariance in unsupervised speech representations
Unsupervised speech representations have taken off, with benchmarks (SUPERB, ZeroSpeech) demonstrating major progress on semi-supervised speech recognition, speech synthesis, and speech-only language modelling. Inspirati…
Language Modellingspeech-recognitionSpeech RecognitionSpeech Synthesis