paper-with-me

홈 › Papers

Deep low-latency joint speech transmission and enhancement over a gaussian channel

2024-04-30 · Mohammad Bokaei, Jesper Jensen, Simon Doclo, Jan Østergaard

Ensuring intelligible speech communication for hearing assistive devices in low-latency scenarios presents significant challenges in terms of speech enhancement, coding and transmission. In this paper, we propose novel solutions for low-latency joint speech transmission and enhancement, leveraging deep neural networks (DNNs). Our approach integrates two state-of-the-art DNN architectures for low-latency speech enhancement and low-latency analog joint source-channel-based transmission, creating a combined low-latency system and jointly training both systems in an end-to-end approach. Due to the computational demands of the enhancement system, this order is suitable when high computational power is unavailable in the decoder, like hearing assistive devices. The proposed system enables the configuration of total latency, achieving high performance even at latencies as low as 3 ms, which is typically challenging to attain. The simulation results provide compelling evidence that a joint enhancement and transmission system is superior to a simple concatenation system in diverse settings, encompassing various wireless channel conditions, latencies, and background noise scenarios.

📄 PDF Abstract BibTeX arXiv:2404.19375

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Enhancement

Similar Papers 제목 키워드 기반

Large Model Empowered Streaming Speech Semantic Communications

2025-01-10 · Zhenzi Weng, Zhijin Qin, Geoffrey Ye Li

In this paper, we introduce a large model-empowered streaming semantic communication system for speech transmission across various languages, named LSSC-ST. Specifically, we devise an edge-device collaborative semantic c…

modelSemantic CommunicationTranslation

SoundStream: An End-to-End Neural Audio Codec

2021-07-07 · Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund 외

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStream relies on a model architecture compose…

CPUDecoderSpeech Enhancementtext-to-speech+1

A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids

2022-10-24 · Abhijeet Bishnu, Ankit Gupta, Mandar Gogate, Kia Dashtipour 외

In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hea…

Lip ReadingSpeech Enhancement

Low-latency Monaural Speech Enhancement with Deep Filter-bank Equalizer

2022-02-14 · Chengshi Zheng, Wenzhe Liu, Andong Li, Yuxuan Ke 외

It is highly desirable that speech enhancement algorithms can achieve good performance while keeping low latency for many applications, such as digital hearing aids, acoustically transparent hearing devices, and public a…

Deep LearningSpeech Enhancement

Large Speech Model Enabled Semantic Communication

2025-12-04 · Yun Tian, Zhijin Qin, Guocheng Lv, Ye Jin 외 arxiv

Existing speech semantic communication systems mainly based on Joint Source-Channel Coding (JSCC) architectures have demonstrated impressive performance, but their effectiveness remains limited by model structures specif…

Semantic Communication