paper-with-me

홈 › Papers

VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space

2024-11-22 · Armani Rodriguez, Silvija Kokalj-Filipovic

Generating high-quality speech efficiently remains a key challenge for generative models in speech synthesis. This paper introduces VQalAttent, a lightweight model designed to generate fake speech with tunable performance and interpretability. Leveraging the AudioMNIST dataset, consisting of human utterances of decimal digits (0-9), our method employs a two-step architecture: first, a scalable vector quantized autoencoder (VQ-VAE) that compresses audio spectrograms into discrete latent representations, and second, a decoder-only transformer that learns the probability model of these latents. Trained transformer generates similar latent sequences, convertible to audio spectrograms by the VQ-VAE decoder, from which we generate fake utterances. Interpreting statistical and perceptual quality of the fakes, depending on the dimension and the extrinsic information of the latent space, enables guided improvements in larger, commercial generative models. As a valuable tool for understanding and refining audio synthesis, our results demonstrate VQalAttent's capacity to generate intelligible speech samples with limited computational resources, while the modularity and transparency of the training pipeline helps easily correlate the analytics with modular modifications, hence providing insights for the more complex models.

📄 PDF Abstract BibTeX arXiv:2411.14642

Code (0)

등록된 구현이 없습니다.

Tasks

Audio SynthesisDecoderSpeech Synthesis

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model

2025-07-07 · Chen Wang, Tianyu Peng, Wen Yang, Yinan Bai 외 arxiv

Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expressive responses. However, the most powe…

From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

2025-09-02 · Nithyashree Sivasubramaniam arxiv

Silent Speech Interfaces (SSIs) have gained attention for their ability to generate intelligible speech from non-acoustic signals. While significant progress has been made in advancing speech generation pipelines, limite…

Speech Recognition

KIT's IWSLT 2020 SLT Translation System

2020-07-01 · WS 2020 7 · Ngoc-Quan Pham, Felix Schneider, Tuan-Nam Nguyen, Thanh-Le Ha 외

This paper describes KIT{'}s submissions to the IWSLT2020 Speech Translation evaluation campaign. We first participate in the simultaneous translation task, in which our simultaneous models are Transformer based and can …

Translation

Scaling Transformers for Low-Bitrate High-Quality Speech Coding

2024-11-29 · Julian D Parker, Anton Smirnov, Jordi Pons, CJ Carr 외

The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models hav…

Quantization

A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis

2025-09-16 · Javeria Amir, Farwa Attaria, Mah Jabeen, Umara Noor 외 arxiv

Recent developments in voice cloning and talking head generation demonstrate impressive capabilities in synthesizing natural speech and realistic lip synchronization. Current methods typically require and are trained on …

Talking Head GenerationText to Speech