paper-with-me

Papers

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

2024-12-03 · Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, Jie Tang

We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, speech rate, and dialect according to user instructions. GLM-4-Voice uses an ultra-low bitrate (175bps), single-codebook speech tokenizer with 12.5Hz frame rate derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. To efficiently transfer knowledge from text to speech modalities, we synthesize speech-text interleaved data from existing text pre-training corpora using a text-to-token model. We continue pre-training from the pre-trained text language model GLM-4-9B with a combination of unsupervised speech data, interleaved speech-text data, and supervised speech-text data, scaling up to 1 trillion tokens, achieving state-of-the-art performance in both speech language modeling and spoken question answering. We then fine-tune the pre-trained model with high-quality conversational speech data, achieving superior performance compared to existing baselines in both conversational ability and speech quality. The open models can be accessed through https://github.com/THUDM/GLM-4-Voice and https://huggingface.co/THUDM/glm-4-voice-9b.

📄 PDF Abstract BibTeX arXiv:2412.02612

Code (1)

thudm/glm-4-voice 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ChatbotLanguage ModelingLanguage ModellingQuestion Answeringspeech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

2025-05-05 · Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang 외

Real-time, intelligent, and natural speech interaction is an essential part of the next-generation human-computer interaction. Recent advancements have showcased the potential of building intelligent spoken chatbots base…

ChatbotDecoderInstruction FollowingQuestion Answering+1

JoyTTS: LLM-based Spoken Chatbot With Voice Cloning

2025-07-03 · Fangru Zhou, Jun Zhao, Guoxin Wang arxiv

JoyTTS is an end-to-end spoken chatbot that combines large language models (LLM) with text-to-speech (TTS) technology, featuring voice cloning capabilities. This project is built upon the open-source MiniCPM-o and CosyVo…

Voice-Based Smart Assistant System for Vehicles using RASA

2023-12-04 · Aditya Paranjape, Yash Patwardhan, Vedant Deshpande, Aniket Darp 외

Conversational AIs, or chatbots, mimic human speech when conversing. Smart assistants facilitate the automation of several tasks that needed human intervention earlier. Because of their accuracy, absence of dependence on…

Discourse-Driven Integrated Dialogue Development Environment for Open-Domain Dialogue Systems

2021-11-01 · CODI 2021 11 · Denis Kuznetsov, Dmitry Evseev, Lidia Ostyakova, Oleg Serikov 외

Development environments for spoken dialogue systems are popular today because they enable rapid creation of the dialogue systems in times when usage of the voice AI Assistants is constantly growing. We describe a graphi…

Spoken Dialogue Systems

SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan

2024-05-08 · You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto 외

The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spok…

DeepFake DetectionFace Swapping