paper-with-me

홈 › Papers

Practical cognitive speech compression

2022-03-08 · Reza Lotfidereshgi, Philippe Gournay

This paper presents a new neural speech compression method that is practical in the sense that it operates at low bitrate, introduces a low latency, is compatible in computational complexity with current mobile devices, and provides a subjective quality that is comparable to that of standard mobile-telephony codecs. Other recently proposed neural vocoders also have the ability to operate at low bitrate. However, they do not produce the same level of subjective quality as standard codecs. On the other hand, standard codecs rely on objective and short-term metrics such as the segmental signal-to-noise ratio that correlate only weakly with perception. Furthermore, standard codecs are less efficient than unsupervised neural networks at capturing speech attributes, especially long-term ones. The proposed method combines a cognitive-coding encoder that extracts an interpretable unsupervised hierarchical representation with a multi stage decoder that has a GAN-based architecture. We observe that this method is very robust to the quantization of representation features. An AB test was conducted on a subset of the Harvard sentences that are commonly used to evaluate standard mobile-telephony codecs. The results show that the proposed method outperforms the standard AMR-WB codec in terms of delay, bitrate and subjective quality.

📄 PDF Abstract BibTeX arXiv:2203.04415

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderQuantization

Similar Papers 제목 키워드 기반

Cognitive Coding of Speech

2021-10-08 · Reza Lotfidereshgi, Philippe Gournay

We propose an approach for cognitive coding of speech by unsupervised extraction of contextual representations in two hierarchical levels of abstraction. Speech attributes such as phoneme identity that last one hundred m…

Dimensionality ReductionQuantization

Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers

2022-11-17 · Tzu-Quan Lin, Tsung-Huan Yang, Chun-Yao Chang, Kuang-Ming Chen 외

Transformer-based self-supervised models have achieved remarkable success in speech processing, but their large size and high inference cost present significant challenges for real-world deployment. While numerous compre…

Knowledge DistillationModel CompressionSelf-Supervised Learning

"Notic My Speech" -- Blending Speech Patterns With Multimedia

2020-06-12 · Dhruva Sahrawat, Yaman Kumar, Shashwat Aggarwal, Yifang Yin 외

Speech as a natural signal is composed of three parts - visemes (visual part of speech), phonemes (spoken part of speech), and language (the imposed structure). However, video as a medium for the delivery of speech and a…

speech-recognitionSpeech RecognitionVideo CompressionVisual Speech Recognition

Research on several key technologies in practical speech emotion recognition

2017-09-27 · Chengwei Huang

In this dissertation the practical speech emotion recognition technology is studied, including several cognitive related emotion types, namely fidgetiness, confidence and tiredness. The high quality of naturalistic emoti…

ClusteringEmotion RecognitionSpeech Emotion Recognition

Towards Audio Codec-based Speech Separation

2024-06-18 · Jia Qi Yip, Shengkui Zhao, Dianwen Ng, Eng Siong Chng 외

Recent improvements in neural audio codec (NAC) models have generated interest in adopting pre-trained codecs for a variety of speech processing applications to take advantage of the efficiencies gained from high compres…

Edge-computingSpeech Separation