paper-with-me

Papers

Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space

2024-09-19 · Sebastião Quintas, Isabelle Ferrané, Thomas Pellegrini

The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text-to-speech systems with voice cloning capabilities, that allow the usage of a larger amount of voices based on short audio segments, it is known that these systems tend to hallucinate and oftentimes produce bad data that will most likely have a negative impact on the downstream task. In the present work, we conduct a set of experiments around zero-shot learning with synthetic speech data for the specific task of speech commands classification. Our results on the Google Speech Commands dataset show that a simple ASR-based filtering method can have a big impact in the quality of the generated data, translating to a better performance. Furthermore, despite the good quality of the generated speech data, we also show that synthetic and real speech can still be easily distinguishable when using self-supervised (WavLM) features, an aspect further explored with a CycleGAN to bridge the gap between the two types of speech material.

📄 PDF Abstract BibTeX arXiv:2409.12745

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionData AugmentationDomain Adaptationspeech-recognitionSpeech Recognitiontext-to-speechText to SpeechVoice CloningZero-Shot Learning

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Batch Normalization 설명 없음
Residual Connection 설명 없음
Tanh Activation 설명 없음
PatchGAN 설명 없음
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Cycle Consistency Loss Cycle Consistency Loss is a type of loss used for generative adversarial networks that performs unpaired image-to-image translation. It was introduced with the…

Similar Papers 제목 키워드 기반

Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU

2020-05-01 · LREC 2020 5 · Thierry Desot, Fran{\c{c}}ois Portet, Michel Vacher

Massive amounts of annotated data greatly contributed to the advance of the machine learning field. However such large data sets are often unavailable for novel tasks performed in realistic environments such as smart hom…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dynamic Time WarpingNatural Language Understanding+6

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech

2025-08-08 · Theresa Pekarek Rosin, Burak Can Kaplan, Stefan Wermter arxiv

Intent recognition (IR) for speech commands is essential for artificial intelligence (AI) assistant systems; however, most existing approaches are limited to short commands and are predominantly developed for English. Th…

Intent Recognition

Speech Command Recognition in Computationally Constrained Environments with a Quadratic Self-organized Operational Layer

2020-11-23 · Mohammad Soltanian, Junaid Malik, Jenni Raitoharju, Alexandros Iosifidis 외

Automatic classification of speech commands has revolutionized human computer interactions in robotic applications. However, employed recognition models usually follow the methodology of deep learning with complicated ne…

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

2025-05-23 · Alan Dao, Dinh Bach Vu, Huy Hoang Ha, Tuan Le Duc Anh 외

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable …

speech-recognitionSpeech Recognitiontext-to-speechText to Speech

PATE-AAE: Incorporating Adversarial Autoencoder into Private Aggregation of Teacher Ensembles for Spoken Command Classification

2021-04-02 · Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chin-Hui Lee

We propose using an adversarial autoencoder (AAE) to replace generative adversarial network (GAN) in the private aggregation of teacher ensembles (PATE), a solution for ensuring differential privacy in speech application…

Generative Adversarial NetworkKeyword SpottingPrivacy Preserving