paper-with-me

Papers

Generative Semantic Communication for Text-to-Speech Synthesis

2024-10-04 · Jiahao Zheng, Jinke Ren, Peng Xu, Zhihao Yuan, Jie Xu, Fangxin Wang, Gui Gui, Shuguang Cui

Semantic communication is a promising technology to improve communication efficiency by transmitting only the semantic information of the source data. However, traditional semantic communication methods primarily focus on data reconstruction tasks, which may not be efficient for emerging generative tasks such as text-to-speech (TTS) synthesis. To address this limitation, this paper develops a novel generative semantic communication framework for TTS synthesis, leveraging generative artificial intelligence technologies. Firstly, we utilize a pre-trained large speech model called WavLM and the residual vector quantization method to construct two semantic knowledge bases (KBs) at the transmitter and receiver, respectively. The KB at the transmitter enables effective semantic extraction, while the KB at the receiver facilitates lifelike speech synthesis. Then, we employ a transformer encoder and a diffusion model to achieve efficient semantic coding without introducing significant communication overhead. Finally, numerical results demonstrate that our framework achieves much higher fidelity for the generated speech than four baselines, in both cases with additive white Gaussian noise channel and Rayleigh fading channel.

📄 PDF Abstract BibTeX arXiv:2410.03459

Code (0)

등록된 구현이 없습니다.

Tasks

QuantizationSemantic CommunicationSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Deep Learning Enabled Semantic Communications with Speech Recognition and Synthesis

2022-05-09 · Zhenzi Weng, Zhijin Qin, Xiaoming Tao, Chengkang Pan 외

In this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication s…

Deep LearningSemantic Communicationspeech-recognitionSpeech Recognition+1

Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks

2023-12-10 · Seo-Hyun Lee, Young-Eun Lee, Soowon Kim, Byung-Kwan Ko 외

Brain-to-speech technology represents a fusion of interdisciplinary applications encompassing fields of artificial intelligence, brain-computer interfaces, and speech synthesis. Neural representation learning based inten…

Representation LearningSpeech Synthesis

Robust Semantic Communications for Speech Transmission

2024-03-08 · Zhenzi Weng, Zhijin Qin

In this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Particularly, we consider the speech-to-text translation (S2TT) …

Generative Adversarial NetworkSemantic CommunicationSpeech-to-TextSpeech-to-Text Translation+1

On the Emotion Understanding of Synthesized Speech

2026-03-17 · Yuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang 외 arxiv

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding r…

Speech Emotion RecognitionSpeech Synthesis

Semantic Co-Speech Gesture Synthesis and Real-Time Control for Humanoid Robots

2025-12-19 · Gang Zhang arxiv

We present an innovative end-to-end framework for synthesizing semantically meaningful co-speech gestures and deploying them in real-time on a humanoid robot. This system addresses the challenge of creating natural, expr…

Gesture Generation