paper-with-me

Papers

Textually Pretrained Speech Language Models

2023-05-22 · NeurIPS 2023 11 · Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Defossez, Gabriel Synnaeve, Emmanuel Dupoux, Roy Schwartz, Yossi Adi

Speech language models (SpeechLMs) process and generate acoustic data only, without textual supervision. In this work, we propose TWIST, a method for training SpeechLMs using a warm-start from a pretrained textual language models. We show using both automatic and human evaluations that TWIST outperforms a cold-start SpeechLM across the board. We empirically analyze the effect of different model design choices such as the speech tokenizer, the pretrained textual model, and the dataset size. We find that model and dataset scale both play an important role in constructing better-performing SpeechLMs. Based on our observations, we present the largest (to the best of our knowledge) SpeechLM both in terms of number of parameters and training data. We additionally introduce two spoken versions of the StoryCloze textual benchmark to further improve model evaluation and advance future research in the field. We make speech samples, code and models publicly available: https://pages.cs.huji.ac.il/adiyoss-lab/twist/ .

📄 PDF Abstract BibTeX arXiv:2305.13009

Code (1)

slp-rl/spokenstorycloze 공식 구현

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

Speakers enhance contextually confusable words

2020-07-01 · ACL 2020 6 · Eric Meinhardt, Eric Bakovic, Leon Bergen

Recent work has found evidence that natural languages are shaped by pressures for efficient communication {---} e.g. the more contextually predictable a word is, the fewer speech sounds or syllables it has (Piantadosi et…

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

2025-02-27 · Weihao wu, Zhiwei Lin, Yixuan Zhou, Jingbei Li 외

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, exist…

DiversityLanguage ModelingLanguage ModellingSpeech Synthesis

Contextually-rich human affect perception using multimodal scene information

2023-03-13 · Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shrikanth Narayanan

The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focuse…

Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

2024-06-12 · Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 외

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

Incorporating Human Explanations for Robust Hate Speech Detection

2024-11-09 · Jennifer L. Chen, Faisal Ladhak, Daniel Li, Noémie Elhadad

Given the black-box nature and complexity of large transformer language models (LM), concerns about generalizability and robustness present ethical implications for domains such as hate speech (HS) detection. Using the c…

Hate Speech Detection