paper-with-me

홈 › Papers

Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation

2019-10-29 · Interspeech 2019 2019 7

We describe Parrotron, an end-to-end-trained speech-to-speech conversion model that maps an input spectrogram directly to another spectrogram, without utilizing any intermediate discrete representation. The network is composed of an encoder, spectrogram and phoneme decoders, followed by a vocoder to synthesize a time-domain waveform. We demonstrate that this model can be trained to normalize speech from any speaker regardless of accent, prosody, and background noise, into the voice of a single canonical target speaker with a fixed accent and consistent articulation and prosody. We further show that this normalization model can be adapted to normalize highly atypical speech from a deaf speaker, resulting in significant improvements in intelligibility and naturalness, measured via a speech recognizer and listening tests. Finally, demonstrating the utility of this model on other speech tasks, we show that the same model architecture can be trained to perform a speech separation task

📄 PDF Abstract BibTeX arXiv:1904.04169

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Streaming Parrotron for on-device speech-to-speech conversion

2022-10-25 · Oleg Rybakov, Fadi Biadsy, Xia Zhang, Liyang Jiang 외

We present a fully on-device streaming Speech2Speech conversion model that normalizes a given input speech directly to synthesized output speech. Deploying such a model on mobile devices pose significant challenges in te…

CPUDecoderQuantizationSTS

Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments

2024-07-07 · Ritabrata Roy Choudhury

The proposed model aims to develop a speech recognition technology for hearing, speech, or cognitively disabled people. All the available technology in the field of speech recognition doesn't come with an interface for c…

speech-recognitionSpeech Recognition

Sign-to-Speech Model for Sign Language Understanding: A Case Study of Nigerian Sign Language

2021-11-01 · Steven Kolawole, Opeyemi Osakuade, Nayan Saxena, Babatunde Kazeem Olorisade

Through this paper, we seek to reduce the communication barrier between the hearing-impaired community and the larger society who are usually not familiar with sign language in the sub-Saharan region of Africa with the l…

object-detectionObject Detection

Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired Listeners

2022-04-08 · Zehai Tu, Ning Ma, Jon Barker

An accurate objective speech intelligibility prediction algorithms is of great interest for many applications such as speech enhancement for hearing aids. Most algorithms measures the signal-to-noise ratios or correlatio…

PredictionSpeech Enhancementspeech-recognitionSpeech Recognition

HASA-net: A non-intrusive hearing-aid speech assessment network

2021-11-10 · Hsin-Tien Chiang, Yi-Chiao Wu, Cheng Yu, Tomoki Toda 외

Without the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models have been applied to build non-intrusive sp…