paper-with-me

홈 › Papers

Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

2024-12-22 · Hankun Wang, Haoran Wang, Yiwei Guo, Zhihan Li, Chenpeng Du, Xie Chen, Kai Yu

Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coherent outputs. There are several potential reasons for this performance degradation: (A) speech tokens mainly provide phonetic information rather than semantic information, (B) the length of speech sequences is much longer than that of text sequences, and (C) paralinguistic information, such as prosody, introduces additional complexity and variability. In this paper, we explore the influence of three key factors separately by transiting the modality from text to speech in an evolving manner. Our findings reveal that the impact of the three factors varies. Factor A has a relatively minor impact, factor B influences syntactical and semantic modeling more obviously, and factor C exerts the most significant impact, particularly in the basic lexical modeling. Based on these findings, we provide insights into the unique challenges of training SLMs and highlight pathways to develop more effective end-to-end SLMs.

📄 PDF Abstract BibTeX arXiv:2412.17048

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain

2025-03-26 · Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao 외

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGe…

Gesture Generation

Variable-rate discrete representation learning

2021-03-10 · Sander Dieleman, Charlie Nash, Jesse Engel, Karen Simonyan

Semantically meaningful information content in perceptual signals is usually unevenly distributed. In speech signals for example, there are often many silences, and the speed of pronunciation can vary considerably. In th…

Representation Learning

AudioLM: a Language Modeling Approach to Audio Generation

2022-09-07 · Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov 외

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this…

Audio GenerationLanguage ModelingLanguage Modelling+1

Text-Free Prosody-Aware Generative Spoken Language Modeling

2021-09-07 · ACL 2022 5 · Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi 외

Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored. Generativ…

Language ModelingLanguage Modelling

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

2024-03-26 · CVPR 2024 1 · Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli 외

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utter…

Gesture Generation