paper-with-me

Papers

JoyTTS: LLM-based Spoken Chatbot With Voice Cloning

2025-07-03 · Fangru Zhou, Jun Zhao, Guoxin Wang arxiv

JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVoice2 models and trained on 2000 hours of conversational data. We have also provided the complete training code to facilitate further development and optimization by the community. On the testing machine seed-tts-zh, it achieves a SS (speaker similarity) score of 0.73 and a WER (Word Error Rate) of 5.09. The code and models, along with training and inference scripts, are available at https://github.com/jdh-algo/JoyTTS.git.

📄 PDF Abstract BibTeX arXiv:2507.02380

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

2026-04-28 · Amanuel Gizachew Abebe, Yasmin Moslem arxiv

Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. I…

Data Augmentation

FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

2026-01-16 · Tanyu Chen, Tairan Chen, Kai Shen, Zhenghua Bao 외 arxiv

Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these models often exhibit limited speaker iden…

Latent linguistic embedding for cross-lingual text-to-speech and voice conversion

2020-10-08 · Hieu-Thi Luong, Junichi Yamagishi

As the recently proposed voice cloning system, NAUTILUS, is capable of cloning unseen voices using untranscribed speech, we investigate the feasibility of using it to develop a unified cross-lingual TTS/VC system. Cross-…

text-to-speechText to SpeechVoice CloningVoice Conversion

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

2024-12-03 · Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang 외

We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ChatbotLanguage Modeling+6

Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech

2024-06-11 · Mateusz Czyżnikiewicz, Łukasz Bondaruk, Jakub Kubiak, Adam Wiącek 외

In this paper we study the impact of augmenting spoken language corpora with domain-specific synthetic samples for the purpose of training a speech recognition system. Using both a conventional neural TTS system and a ze…

speech-recognitionSpeech RecognitionVoice Cloning