paper-with-me

홈 › Papers

Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

2023-09-14 · Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, Shinji Watanabe

We propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens from self-supervised speech features and uses special tokens to enable multitask learning. Compared to a single-task model, VoxtLM exhibits a significant improvement in speech synthesis, with improvements in both speech intelligibility from 28.9 to 5.6 and objective quality from 2.68 to 3.90. VoxtLM also improves speech generation and speech recognition performance over the single-task counterpart. Further, VoxtLM is trained with publicly available data and training recipes and model checkpoints are open-sourced to make fully reproducible work.

📄 PDF Abstract BibTeX arXiv:2309.07937

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech SynthesisText Generation

Similar Papers 제목 키워드 기반

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

2025-10-23 · Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang 외 arxiv

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying speech enhancement (SE) tasks remains u…

Reinforcement LearningSpeech EnhancementSpeech Separation

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

2021-10-14 · ACL 2022 5 · Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang 외

Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderQuantization+7

ToXCL: A Unified Framework for Toxic Speech Detection and Explanation

2024-03-25 · Nhat M. Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu 외

The proliferation of online toxic speech is a pertinent problem posing threats to demographic groups. While explicit toxic speech contains offensive lexical signals, implicit one consists of coded or indirect language. T…

DecoderKnowledge DistillationText Generation

Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR

2023-11-08 · Qian Chen, Wen Wang, Qinglin Zhang, Siqi Zheng 외

Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and …

Decoder

OpusLM: A Family of Open Unified Speech Language Models

2025-06-21 · Jinchuan Tian, William Chen, Yifan Peng, Jiatong Shi 외

This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text language models, the OpusLMs are continuous…

Decoderspeech-recognitionSpeech RecognitionSpeech Synthesis