Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
We propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens from self-supervised speech features and uses special tokens to enable multitask learning. Compared to a single-task model, VoxtLM exhibits a significant improvement in speech synthesis, with improvements in both speech intelligibility from 28.9 to 5.6 and objective quality from 2.68 to 3.90. VoxtLM also improves speech generation and speech recognition performance over the single-task counterpart. Further, VoxtLM is trained with publicly available data and training recipes and model checkpoints are open-sourced to make fully reproducible work.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech SynthesisText GenerationSimilar Papers 제목 키워드 기반
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement
Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying speech enhancement (SE) tasks remains u…
Reinforcement LearningSpeech EnhancementSpeech SeparationSpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderQuantization+7ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
The proliferation of online toxic speech is a pertinent problem posing threats to demographic groups. While explicit toxic speech contains offensive lexical signals, implicit one consists of coded or indirect language. T…
DecoderKnowledge DistillationText GenerationLoss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and …
DecoderOpusLM: A Family of Open Unified Speech Language Models
This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text language models, the OpusLMs are continuous…
Decoderspeech-recognitionSpeech RecognitionSpeech Synthesis