Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example
Speech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes.
Code (1)
Tasks
Voice ConversionSimilar Papers 제목 키워드 기반
Pathological voice adaptation with autoencoder-based voice conversion
In this paper, we propose a new approach to pathological speech synthesis. Instead of using healthy speech as a source, we customise an existing pathological speech sample to a new speaker's voice characteristics. This a…
Speech SynthesisVoice ConversionAn Objective Evaluation Framework for Pathological Speech Synthesis
The development of pathological speech systems is currently hindered by the lack of a standardised objective evaluation framework. In this work, (1) we utilise existing detection and analysis techniques to propose a gene…
Speech SynthesisVoice ConversionSelf-Supervised Speech Representations Preserve Speech Characteristics while Anonymizing Voices
Collecting speech data is an important step in training speech recognition systems and other speech-based machine learning models. However, the issue of privacy protection is an increasing concern that must be addressed.…
Speaker Verificationspeech-recognitionSpeech RecognitionVoice ConversionVOTE400(Voide Of The Elderly 400 Hours): A Speech Dataset to Study Voice Interface for Elderly-Care
This paper introduces a large-scale Korean speech dataset, called VOTE400, that can be used for analyzing and recognizing voices of the elderly people. The dataset includes about 300 hours of continuous dialog speech and…
speech-recognitionSpeech RecognitionTowards Identity Preserving Normal to Dysarthric Voice Conversion
We present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation o…
Data AugmentationDecision Makingspeech-recognitionSpeech Recognition+1