paper-with-me

홈 › Papers

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

2024-11-13 · Dingdong Wang, Mingyu Cui, Dongchao Yang, Xueyuan Chen, Helen Meng

With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamlessly. Compared to most studies that focus on continuous speech features, although discrete-token based LLMs have shown promising results on certain tasks, the performance gap between these two paradigms is rarely explored. In this paper, we present a fair and thorough comparison between discrete and continuous features across a variety of semantic-related tasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal that continuous features generally outperform discrete tokens, particularly in tasks requiring fine-grained semantic understanding. Moreover, this study goes beyond surface-level comparison by identifying key factors behind the under-performance of discrete tokens, such as limited token granularity and inefficient information retention. To enhance the performance of discrete tokens, we explore potential aspects based on our analysis. We hope our results can offer new insights into the opportunities for advancing discrete speech tokens in Speech LLMs.

📄 PDF Abstract BibTeX arXiv:2411.08742

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Continuous Speech Synthesis using per-token Latent Diffusion

2024-10-21 · Arnon Turetzky, Nimrod Shabtay, Slava Shechtman, Hagai Aronowitz 외

The success of autoregressive transformer models with discrete tokens has inspired quantization-based approaches for continuous modalities, though these often limit reconstruction quality. We therefore introduce SALAD, a…

Image GenerationQuantizationSpeech Synthesistext-to-speech+1

Analysing the Language of Neural Audio Codecs

2025-09-01 · Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando 외 arxiv

This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to li…

Speech Recognition

Benchmarking Prosody Encoding in Discrete Speech Tokens

2025-08-15 · Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu arxiv

Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for va…

Self-Supervised Learning

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

2025-08-25 · Dingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang 외 arxiv

With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio…

Spoken Language UnderstandingSelf-Supervised Learning

DASB -- Discrete Audio and Speech Benchmark

2024-06-20 · Pooneh Mousavi, Luca Della Libera, Jarod Duret, Artem Ploujnikov 외

Discrete audio tokens have recently gained considerable attention for their potential to connect audio and language processing, enabling the creation of modern multimodal large language models. Ideal audio tokens must ef…

BenchmarkingEmotion Recognitionintent-classificationIntent Classification+7