paper-with-me

홈 › Papers

Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space

2025-05-19 · Zhengrui Ma, Yang Feng, Chenze Shao, Fandong Meng, Jie zhou, Min Zhang

We introduce SLED, an alternative approach to speech language modeling by encoding speech waveforms into sequences of continuous latent representations and modeling them autoregressively using an energy distance objective. The energy distance offers an analytical measure of the distributional gap by contrasting simulated and target samples, enabling efficient training to capture the underlying continuous autoregressive distribution. By bypassing reliance on residual vector quantization, SLED avoids discretization errors and eliminates the need for the complicated hierarchical architectures common in existing speech language models. It simplifies the overall modeling pipeline while preserving the richness of speech information and maintaining inference efficiency. Empirical results demonstrate that SLED achieves strong performance in both zero-shot and streaming speech synthesis, showing its potential for broader applications in general-purpose speech language models.

📄 PDF Abstract BibTeX arXiv:2505.13181

Code (1)

ictnlp/sled-tts 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingQuantizationSpeech Synthesis

Similar Papers 제목 키워드 기반

A Spectral Energy Distance for Parallel Speech Synthesis

2020-08-03 · NeurIPS 2020 12 · Alexey A. Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek 외

Speech synthesis is an important practical generative modeling problem that has seen great progress over the last few years, with likelihood-based autoregressive neural models now outperforming traditional concatenative …

scoring ruleSpeech Synthesis

Learning Approximate Inference Networks for Structured Prediction

2018-03-09 · ICLR 2018 1 · Lifu Tu, Kevin Gimpel

Structured prediction energy networks (SPENs; Belanger & McCallum 2016) use neural network architectures to define energy functions that can capture arbitrary dependencies among parts of structured outputs. Prior work us…

Language ModelingLanguage ModellingMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+3

Are discrete units necessary for Spoken Language Modeling?

2022-03-11 · Tu Anh Nguyen, Benoit Sagot, Emmanuel Dupoux

Recent work in spoken language modeling shows the possibility of learning a language unsupervisedly from raw audio without any text labels. The approach relies first on transforming the audio into a sequence of discrete …

Language ModelingLanguage Modelling

First Automatic Fongbe Continuous Speech Recognition System: Development of Acoustic Models and Language Models

2017-01-21 · Fréjus Laleye, Laurent Besacier, Eugène Ezin, Cina Motamed.

This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe). The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe. The problem …

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

2025-02-03 · Weiwei Lin, Chenghan He

We propose a novel autoregressive modeling approach for speech synthesis, combining a variational autoencoder (VAE) with a multi-modal latent space and an autoregressive model that uses Gaussian Mixture Models (GMM) as t…

QuantizationSpeech Synthesis