paper-with-me

Papers

Comparing Discrete and Continuous Space LLMs for Speech Recognition

2024-09-01 · Yaoxun Xu, Shi-Xiong Zhang, Jianwei Yu, Zhiyong Wu, Dong Yu

This paper investigates discrete and continuous speech representations in Large Language Model (LLM)-based Automatic Speech Recognition (ASR), organizing them by feature continuity and training approach into four categories: supervised and unsupervised for both discrete and continuous types. We further classify LLMs based on their input and autoregressive feedback into continuous and discrete-space models. Using specialized encoders and comparative analysis with a Joint-Training-From-Scratch Language Model (JTFS LM) and pre-trained LLaMA2-7b, we provide a detailed examination of their effectiveness. Our work marks the first extensive comparison of speech representations in LLM-based ASR and explores various modeling techniques. We present an open-sourced achievement of a state-of-the-art Word Error Rate (WER) of 1.69\% on LibriSpeech using a HuBERT encoder, offering valuable insights for advancing ASR and natural language processing (NLP) research.

📄 PDF Abstract BibTeX arXiv:2409.00800

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

2025-08-25 · Dingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang 외 arxiv

With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio…

Spoken Language UnderstandingSelf-Supervised Learning

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

2024-11-13 · Dingdong Wang, Mingyu Cui, Dongchao Yang, Xueyuan Chen 외

With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamlessly. Compared to most studies that focus…

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

2026-06-08 · Ming-Hao Hsu, Yuxuan Hu, Shujie Liu, Jinyu Li 외 arxiv

Large language models (LLMs) provide a powerful reasoning backbone for speech understanding, but integrating continuous acoustic signals into a frozen LLM remains challenging. Existing speech-to-LLM interfaces typically …

Emotion RecognitionSpeech Recognition

DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation

2023-10-26 · Yongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye 외

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial probl…

Image GenerationSpeech-to-Speech TranslationTranslation

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

2023-10-07 · Zhihao Du, JiaMing Wang, Qian Chen, Yunfei Chu 외

Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for audio-and-text large language models (LLM…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+14