paper-with-me

홈 › Papers

On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models

2026-06-22 · Shunsuke Kando, Wataru Nakata, Shinnosuke Takamichi, Yusuke Miyao arxiv

Generative Spoken Language Modeling (GSLM) enables text-free speech modeling by training language models (LMs) using discrete speech representations instead of textual transcription. In this paper, we investigate the performance of GSLM on speech synthesis and continuation using discrete speech representations with varying bitrates. We segment speech representations with fixed widths and train K-means models in multiple cluster sizes, resulting in various bitrate settings. We demonstrate that intelligible and natural speech can be synthesized at lower bitrate settings than the baseline. Furthermore, speech continuation quality remains stable at lower bitrates across multiple metrics, suggesting that the conventional GSLM setting may be redundant for effective speech generation. Although LLM-based metrics show higher correlation with human subjective score than conventional metrics, it remains low, highlighting the need for more stable automatic evaluation methods.

📄 PDF Abstract BibTeX arXiv:2606.23285

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

2025-05-23 · Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi

The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, the…

Speech TokenizationSpoken Language Understanding

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

2025-06-04 · Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain 외

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22…

Speech Synthesistext-to-speechText to Speech

Nonparametric Regression under Cluster Sampling

2024-03-07 · Yuya Shimizu

This paper develops a general asymptotic theory for nonparametric kernel regression in the presence of cluster dependence. We examine nonparametric density estimation, Nadaraya-Watson kernel regression, and local linear …

Density Estimationregressionvalid

Secrets of GrabCut and Kernel K-Means

2015-12-01 · ICCV 2015 12 · Meng Tang, Ismail Ben Ayed, Dmitrii Marin, Yuri Boykov

The log-likelihood energy term in popular model-fitting segmentation methods, e.g. Zhu&Yuille, Chan-Vese, GrabCut, is presented as a generalized "probabilistic K-means" energy for color space clustering. This interpretat…

ClusteringSegmentation

The spatial scale dimension of speech processing in the human brain

2022-01-19 · Philipp Kellmeyer, Roland Berkemeier, Tonio Ball

In the past three decades, neuroimaging has provided important insights into structure-function relationships in the human brain. Recently, however, the methods for analyzing functional magnetic resonance imaging (fMRI) …