paper-with-me

홈 › Papers

Vision-Integrated High-Quality Neural Speech Coding

2025-05-29 · Yao Guo, Yang Ai, Rui-Chen Zheng, Hui-Peng Du, Xiao-Hang Jiang, Zhen-Hua Ling

This paper proposes a novel vision-integrated neural speech codec (VNSC), which aims to enhance speech coding quality by leveraging visual modality information. In VNSC, the image analysis-synthesis module extracts visual features from lip images, while the feature fusion module facilitates interaction between the image analysis-synthesis module and the speech coding module, transmitting visual information to assist the speech coding process. Depending on whether visual information is available during the inference stage, the feature fusion module integrates visual features into the speech coding module using either explicit integration or implicit distillation strategies. Experimental results confirm that integrating visual information effectively improves the quality of the decoded speech and enhances the noise robustness of the neural speech codec, without increasing the bitrate.

📄 PDF Abstract BibTeX arXiv:2505.23379

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LACE: A light-weight, causal model for enhancing coded speech through adaptive convolutions

2023-07-13 · Jan Büthe, Jean-Marc Valin, Ahmed Mustafa

Classical speech coding uses low-complexity postfilters with zero lookahead to enhance the quality of coded speech, but their effectiveness is limited by their simplicity. Deep Neural Networks (DNNs) can be much more eff…

RedApt: An Adaptor for wav2vec 2 Encoding \\ Faster and Smaller Speech Translation without Quality Compromise

2022-10-16 · Jinming Zhao, Hao Yang, Gholamreza Haffari, Ehsan Shareghi

Pre-trained speech Transformers in speech translation (ST) have facilitated state-of-the-art (SotA) results; yet, using such encoders is computationally expensive. To improve this, we present a novel Reducer Adaptor bloc…

Translation

Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder

2019-10-14 · Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia S. C. Lim 외

In order to efficiently transmit and store speech signals, speech codecs create a minimally redundant representation of the input signal which is then decoded at the receiver with the best possible perceptual quality. In…

Decoder

ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation

2026-01-16 · Zhuoyue Gao, Xiaohui Wang, Xiaocui Yang, Wen Zhang 외 arxiv

Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing spee…

Empathetic Response GenerationSpeech Synthesis

XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System

2020-06-11 · Peiling Lu, Jie Wu, Jian Luan, Xu Tan 외

This paper presents XiaoiceSing, a high-quality singing voice synthesis system which employs an integrated network for spectrum, F0 and duration modeling. We follow the main architecture of FastSpeech while proposing som…

RhythmSinging Voice SynthesisVocal Bursts Intensity Prediction