paper-with-me

Papers

DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion

2024-06-12 · Ziqian Ning, Shuai Wang, Pengcheng Zhu, Zhichao Wang, Jixun Yao, Lei Xie, Mengxiao Bi

Streaming voice conversion has become increasingly popular for its potential in real-time applications. The recently proposed DualVC 2 has achieved robust and high-quality streaming voice conversion with a latency of about 180ms. Nonetheless, the recognition-synthesis framework hinders end-to-end optimization, and the instability of automatic speech recognition (ASR) model with short chunks makes it challenging to further reduce latency. To address these issues, we propose an end-to-end model, DualVC 3. With speaker-independent semantic tokens to guide the training of the content encoder, the dependency on ASR is removed and the model can operate under extremely small chunks, with cascading errors eliminated. A language model is trained on the content encoder output to produce pseudo context by iteratively predicting future frames, providing more contextual information for the decoder to improve conversion quality. Experimental results demonstrate that DualVC 3 achieves comparable performance to DualVC 2 in subjective and objective metrics, with a latency of only 50 ms.

📄 PDF Abstract BibTeX arXiv:2406.07846

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion

2023-09-27 · Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Shuai Wang 외

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective throu…

DecoderKnowledge DistillationVoice Conversion

DualVC: Dual-mode Voice Conversion using Intra-model Knowledge Distillation and Hybrid Predictive Coding

2023-05-21 · Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Jixun Yao 외

Voice conversion is an increasingly popular technology, and the growing number of real-time applications requires models with streaming conversion capabilities. Unlike typical (non-streaming) voice conversion, which can …

Data AugmentationDecoderKnowledge DistillationVoice Conversion

Learning Multiple Object States from Actions via Large Language Models

2024-05-02 · Masatoshi Tateno, Takuma Yagi, Ryosuke Furuta, Yoichi Sato

Recognizing the states of objects in a video is crucial in understanding the scene beyond actions and objects. For instance, an egg can be raw, cracked, and whisked while cooking an omelet, and these states can coexist s…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONObjectWorld Knowledge

Exploring the Best Practices of Query Expansion with Large Language Models

2024-01-12 · Le Zhang, Yihong Wu, Qian Yang, Jian-Yun Nie

Large Language Models (LLMs) are foundational in language technologies, particularly in information retrieval (IR). Previous studies have utilized LLMs for query expansion, achieving notable improvements in IR. In this p…

Information RetrievalRe-RankingRetrievalText Generation+1

Exploiting Pseudo Future Contexts for Emotion Recognition in Conversations

2023-06-27 · Yinyi Wei, Shuaipeng Liu, Hailei Yan, Wei Ye 외

With the extensive accumulation of conversational data on the Internet, emotion recognition in conversations (ERC) has received increasing attention. Previous efforts of this task mainly focus on leveraging contextual an…

Emotion Recognition