paper-with-me

홈 › Papers

Long-Form Speech Generation with Spoken Language Models

2024-12-24 · Se Jin Park, Julian Salazar, Aren Jansen, Keisuke Kinoshita, Yong Man Ro, RJ Skerry-Ryan

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, current spoken language models struggle to generate plausible speech past tens of seconds, from high temporal resolution of speech tokens causing loss of coherence, to architectural issues with long-sequence training or extrapolation, to memory costs at inference time. With these considerations we propose SpeechSSM, the first speech language model to learn from and sample long-form spoken audio (e.g., 16 minutes of read or extemporaneous speech) in a single decoding session without text intermediates, based on recent advances in linear-time sequence modeling. Furthermore, to address growing challenges in spoken language evaluation, especially in this new long-form setting, we propose: new embedding-based and LLM-judged metrics; quality measurements over length and time; and a new benchmark for long-form speech processing and generation, LibriSpeech-Long. Speech samples and the dataset are released at https://google.github.io/tacotron/publications/speechssm/

📄 PDF Abstract BibTeX arXiv:2412.18603

Code (1)

google-deepmind/librispeech-long 공식 구현

Tasks

FormLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

2024-06-18 · Kentaro Mitsui, Koh Mitsuda, Toshiaki Wakatsuki, Yukiya Hono 외

Multimodal language models that process both text and speech have a potential for applications in spoken dialogue systems. However, current models face two major challenges in response generation latency: (1) generating …

Language ModelingLanguage ModellingQuestion AnsweringResponse Generation+1

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering

2025-11-12 · Jiliang Hu, Zuchao Li, Baoyuan Qi, Liu Guoming 외 arxiv

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of ret…

Cross-Modal RetrievalSpeech RecognitionQuestion AnsweringText Retrieval

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

2025-07-21 · Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin 외 arxiv

Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking process before responding. In contrast, huma…

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

2025-01-01 · Haitian Lu, Gaofeng Cheng, Liuping Luo, Leying Zhang 외

Recently, ``textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack se…

Dialogue GenerationLanguage ModelingLanguage Modelling

DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion

2026-01-30 · Yuxuan Lou, Ziming Wu, Yaochen Wang, Yong Liu 외 arxiv

Current speech language models generate responses directly without explicit reasoning, leading to errors that cannot be corrected once audio is produced. We introduce \textbf{``Silent Thought, Spoken Answer''} -- a parad…