Deep Learning Enabled Semantic Communications with Speech Recognition and Synthesis
In this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication system, respectively. First, the speech recognition-related semantic features are extracted for transmission by a joint semantic-channel encoder and the text is recovered at the receiver based on the received semantic features, which significantly reduces the required amount of data transmission without performance degradation. Then, we perform speech synthesis at the receiver, which dedicates to re-generate the speech signals by feeding the recognized text and the speaker information into a neural network module. To enable the DeepSC-ST adaptive to dynamic channel environments, we identify a robust model to cope with different channel conditions. According to the simulation results, the proposed DeepSC-ST significantly outperforms conventional communication systems and existing DL-enabled communication systems, especially in the low signal-to-noise ratio (SNR) regime. A software demonstration is further developed as a proof-of-concept of the DeepSC-ST.
Code (1)
Tasks
Deep LearningSemantic Communicationspeech-recognitionSpeech RecognitionSpeech SynthesisSimilar Papers 제목 키워드 기반
Semantic Communications for Speech Recognition
The traditional communications transmit all the source data represented by bits, regardless of the content of source and the semantic information required by the receiver. However, in some applications, the receiver only…
Semantic Communicationspeech-recognitionSpeech RecognitionSemantic Communication Systems for Speech Transmission
Semantic communications could improve the transmission efficiency significantly by exploring the semantic information. In this paper, we make an effort to recover the transmitted speech signals in the semantic communicat…
Semantic CommunicationRobust Semantic Communications for Speech Transmission
In this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Particularly, we consider the speech-to-text translation (S2TT) …
Generative Adversarial NetworkSemantic CommunicationSpeech-to-TextSpeech-to-Text Translation+1CASSANDRA: A multipurpose configurable voice-enabled human-computer-interface
Voice enabled human computer interfaces (HCI) that integrate automatic speech recognition, text-to-speech synthesis and natural language understanding have become a commodity, introduced by the immersion of smart phones …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Natural Language UnderstandingQuestion Answering+6Speech Synthesis for Low Resource Languages using Transliteration Enabled Transfer Learning
In the area of Human Computer Interaction (HCI), Text To Speech (TTS) synthesis has received a significant boost in recent years, especially with the development of various deep learning techniques capable of generating …
speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+3