Textually Pretrained Speech Language Models
Speech language models (SpeechLMs) process and generate acoustic data only, without textual supervision. In this work, we propose TWIST, a method for training SpeechLMs using a warm-start from a pretrained textual language models. We show using both automatic and human evaluations that TWIST outperforms a cold-start SpeechLM across the board. We empirically analyze the effect of different model design choices such as the speech tokenizer, the pretrained textual model, and the dataset size. We find that model and dataset scale both play an important role in constructing better-performing SpeechLMs. Based on our observations, we present the largest (to the best of our knowledge) SpeechLM both in terms of number of parameters and training data. We additionally introduce two spoken versions of the StoryCloze textual benchmark to further improve model evaluation and advance future research in the field. We make speech samples, code and models publicly available: https://pages.cs.huji.ac.il/adiyoss-lab/twist/ .
Code (1)
Tasks
Language ModellingSimilar Papers 제목 키워드 기반
Speakers enhance contextually confusable words
Recent work has found evidence that natural languages are shaped by pressures for efficient communication {---} e.g. the more contextually predictable a word is, the fewer speech sounds or syllables it has (Piantadosi et…
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, exist…
DiversityLanguage ModelingLanguage ModellingSpeech SynthesisContextually-rich human affect perception using multimodal scene information
The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focuse…
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…
ChatbotLanguage ModelingLanguage ModellingLarge Language ModelIncorporating Human Explanations for Robust Hate Speech Detection
Given the black-box nature and complexity of large transformer language models (LM), concerns about generalizability and robustness present ethical implications for domains such as hate speech (HS) detection. Using the c…
Hate Speech Detection