paper-with-me

Papers

Large Speech Model Enabled Semantic Communication

2025-12-04 · Yun Tian, Zhijin Qin, Guocheng Lv, Ye Jin, Kaibin Huang, Zhu Han arxiv

Existing speech semantic communication systems mainly based on Joint Source-Channel Coding (JSCC) architectures have demonstrated impressive performance, but their effectiveness remains limited by model structures specifically designed for particular tasks and datasets. Recent advances indicate that generative large models pre-trained on massive datasets, can achieve outstanding performance arexhibit exceptional performance across diverse downstream tasks with minimal fine-tuning. To exploit the rich semantic knowledge embedded in large models and enable adaptive transmission over lossy channels, we propose a Large Speech Model enabled Semantic Communication (LargeSC) system. Simultaneously achieving adaptive compression and robust transmission over lossy channels remains challenging, requiring trade-offs among compression efficiency, speech quality, and latency. In this work, we employ the Mimi as a speech codec, converting speech into discrete tokens compatible with existing network architectures. We propose an adaptive controller module that enables adaptive transmission and in-band Unequal Error Protection (UEP), dynamically adjusting to both speech content and packet loss probability under bandwidth constraints. Additionally, we employ Low-Rank Adaptation (LoRA) to finetune the Moshi foundation model for generative recovery of lost speech tokens. Simulation results show that the proposed system supports bandwidths ranging from 550 bps to 2.06 kbps, outperforms conventional baselines in speech quality under high packet loss rates and achieves an end-to-end latency of approximately 460 ms, thereby demonstrating its potential for real-time deployment.

📄 PDF Abstract BibTeX arXiv:2512.04711

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Communication

Similar Papers 제목 키워드 기반

Deep Learning Enabled Semantic Communications with Speech Recognition and Synthesis

2022-05-09 · Zhenzi Weng, Zhijin Qin, Xiaoming Tao, Chengkang Pan 외

In this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication s…

Deep LearningSemantic Communicationspeech-recognitionSpeech Recognition+1

Semantic MIMO Systems for Speech-to-Text Transmission

2024-05-13 · Zhenzi Weng, Zhijin Qin, Huiqiang Xie, Xiaoming Tao 외

Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission…

Semantic CommunicationSpeech-to-Text

Semantic Communication Systems for Speech Transmission

2021-02-24 · Zhenzi Weng, Zhijin Qin

Semantic communications could improve the transmission efficiency significantly by exploring the semantic information. In this paper, we make an effort to recover the transmitted speech signals in the semantic communicat…

Semantic Communication

Robust Semantic Communications for Speech Transmission

2024-03-08 · Zhenzi Weng, Zhijin Qin

In this paper, we propose a robust semantic communication system for speech transmission, named Ross-S2T, by delivering the essential semantic information. Particularly, we consider the speech-to-text translation (S2TT) …

Generative Adversarial NetworkSemantic CommunicationSpeech-to-TextSpeech-to-Text Translation+1

Semantic Communications for Speech Recognition

2021-07-22 · Zhenzi Weng, Zhijin Qin, Geoffrey Ye Li

The traditional communications transmit all the source data represented by bits, regardless of the content of source and the semantic information required by the receiver. However, in some applications, the receiver only…

Semantic Communicationspeech-recognitionSpeech Recognition