Enhanced Voice Post Processing Using Voice Decoder Guidance Indicators
Voice enhancement and voice coding are imperative and important functions in a voice-communication system. However, both functions are commonly treated independently, even though both utilize similar features of the underlying signals. Our proposal is to leverage information from one function to the benefit of the other. Specifically, our proposed changes are focused on changes to the voice enhancement at the downlink side and utilizing information of the voice decoding. Preliminary results show that such an approach results in improved quality. Additionally, suggestions are provided on future extensions of the proposed concept.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderSimilar Papers 제목 키워드 기반
An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in o…
DecoderSinging Voice SynthesisEvaluation of Google's Voice Recognition and Sentence Classification for Health Care Applications
This study examined the use of voice recognition technology in perioperative services (Periop) to enable Periop staff to record workflow milestones using mobile technology. The use of mobile technology to improve patient…
SentenceSentence ClassificationReal-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis
Singing voice conversion is to convert the source singing voice into the target singing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effective…
AttributeDecoderVoice ConversionVoice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module
State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech q…
Speech Synthesistext-to-speechText to SpeechVoice ConversionMinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models f…
Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+4