paper-with-me

홈 › Papers

Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis

2024-12-02 · David Gimeno-Gómez, Catarina Botelho, Anna Pompili, Alberto Abad, Carlos-D. Martínez-Hinarejos

Recent works in pathological speech analysis have increasingly relied on powerful self-supervised speech representations, leading to promising results. However, the complex, black-box nature of these embeddings and the limited research on their interpretability significantly restrict their adoption for clinical diagnosis. To address this gap, we propose a novel, interpretable framework specifically designed to support Parkinson's Disease (PD) diagnosis. Through the design of simple yet effective cross-attention mechanisms for both embedding- and temporal-level analysis, the proposed framework offers interpretability from two distinct but complementary perspectives. Experimental findings across five well-established speech benchmarks for PD detection demonstrate the framework's capability to identify meaningful speech patterns within self-supervised representations for a wide range of assessment tasks. Fine-grained temporal analyses further underscore its potential to enhance the interpretability of deep-learning pathological speech models, paving the way for the development of more transparent, trustworthy, and clinically applicable computer-assisted diagnosis systems in this domain. Moreover, in terms of classification accuracy, our method achieves results competitive with state-of-the-art approaches, while also demonstrating robustness in cross-lingual scenarios when applied to spontaneous speech production.

📄 PDF Abstract BibTeX arXiv:2412.02006

Code (1)

david-gimeno/interpreting-ssl-parkinson-speech 공식 구현

Similar Papers 제목 키워드 기반

ULTra: Unveiling Latent Token Interpretability in Transformer Based Understanding

2024-11-15 · Hesam Hosseini, Ghazal Hosseini Mighan, Amirabbas Afzali, Sajjad Amini 외

Transformers have revolutionized Computer Vision (CV) and Natural Language Processing (NLP) through self-attention mechanisms. However, due to their complexity, their latent token representations are often difficult to i…

SegmentationSemantic SegmentationText SummarizationUnsupervised Semantic Segmentation

Discrete Speech Unit Extraction via Independent Component Analysis

2025-01-11 · Tomohiko Nakamura, Kwanghee Choi, Keigo Hojo, Yoshiaki Bando 외

Self-supervised speech models (S3Ms) have become a common tool for the speech processing community, leveraging representations for downstream tasks. Clustering S3M representations yields discrete speech units (DSUs), whi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeech-recognition+1

Revisiting transposed convolutions for interpreting raw waveform sound event recognition CNNs by sonification

2021-09-29 · Sarthak Yadav, Mary Ellen Foster

The majority of recent work on the interpretability of audio and speech processing deep neural networks (DNNs) interprets spectral information modelled by the first layer, relying solely on visual means of interpretation…

SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations

2024-03-10 · Amit Meghanani, Thomas Hain

There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used…

Automatic Speech RecognitionData AugmentationGPUPhoneme Recognition+3

Evaluating context-invariance in unsupervised speech representations

2022-10-27 · Mark Hallap, Emmanuel Dupoux, Ewan Dunbar

Unsupervised speech representations have taken off, with benchmarks (SUPERB, ZeroSpeech) demonstrating major progress on semi-supervised speech recognition, speech synthesis, and speech-only language modelling. Inspirati…

Language Modellingspeech-recognitionSpeech RecognitionSpeech Synthesis