Sotto Voce: Federated Speech Recognition with Differential Privacy Guarantees
Speech data is expensive to collect, and incredibly sensitive to its sources. It is often the case that organizations independently collect small datasets for their own use, but often these are not performant for the demands of machine learning. Organizations could pool these datasets together and jointly build a strong ASR system; sharing data in the clear, however, comes with tremendous risk, in terms of intellectual property loss as well as loss of privacy of the individuals who exist in the dataset. In this paper, we offer a potential solution for learning an ML model across multiple organizations where we can provide mathematical guarantees limiting privacy loss. We use a Federated Learning approach built on a strong foundation of Differential Privacy techniques. We apply these to a senone classification prototype and demonstrate that the model improves with the addition of private data while still respecting privacy.
Code (0)
등록된 구현이 없습니다.
Tasks
Federated Learningspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks
The availability of digital devices operated by voice is expanding rapidly. However, the applications of voice interfaces are still restricted. For example, speaking in public places becomes an annoyance to the surroundi…
speech-recognitionSpeech RecognitionNasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-…
LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders
Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…
Speech EnhancementSpeech SynthesisRT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every comp…
Speech EnhancementVOCE Corpus: Ecologically Collected Speech Annotated with Physiological and Psychological Stress Assessments
Public speaking is a widely requested professional skill, and at the same time an activity that causes one of the most common adult phobias (Miller and Stone, 2009). It is also known that the study of stress under labora…
Emotion Recognition