Deep Transform: Time-Domain Audio Error Correction via Probabilistic Re-Synthesis
In the process of recording, storage and transmission of time-domain audio signals, errors may be introduced that are difficult to correct in an unsupervised way. Here, we train a convolutional deep neural network to re-synthesize input time-domain speech signals at its output layer. We then use this abstract transformation, which we call a deep transform (DT), to perform probabilistic re-synthesis on further speech (of the same speaker) which has been degraded. Using the convolutive DT, we demonstrate the recovery of speech audio that has been subject to extreme degradation. This approach may be useful for correction of errors in communications devices.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unsupervised domain adaptation for speech recognition with unsupervised error correction
The transcription quality of automatic speech recognition (ASR) systems degrades significantly when transcribing audios coming from unseen domains. We propose an unsupervised error correction method for unsupervised ASR …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDomain Adaptation+3Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition
Language models (LMs) have long been used to improve results of automatic speech recognition (ASR) systems, but they are unaware of the errors that ASR systems make. Error correction models are designed to fix ASR errors…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Denoisingspeech-recognition+3Hybrid phonetic-neural model for correction in speech recognition systems
Automatic speech recognition (ASR) is a relevant area in multiple settings because it provides a natural communication mechanism between applications and users. ASRs often fail in environments that use language specific …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep Learningspeech-recognition+1Real-time error correction and performance aid for MIDI instruments
Making a slight mistake during live music performance can easily be spotted by an astute listener, even if the performance is an improvisation or an unfamiliar piece. An example might be a highly dissonant chord played b…
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
Speaker Diarization (SD) systems are typically audio-based and operate independently of the ASR system in traditional speech transcription pipelines and can have speaker errors due to SD and/or ASR reconciliation, especi…
Language ModelingLanguage Modellingspeaker-diarizationSpeaker Diarization