Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
Acoustic word embeddings (AWEs) are vector representations of spoken words. An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE). In the past, the CAE method has been associated with traditional MFCC features. Representations obtained from self-supervised learning (SSL)-based speech models such as HuBERT, Wav2vec2, etc., are outperforming MFCC in many downstream tasks. However, they have not been well studied in the context of learning AWEs. This work explores the effectiveness of CAE with SSL-based speech representations to obtain improved AWEs. Additionally, the capabilities of SSL-based speech models are explored in cross-lingual scenarios for obtaining AWEs. Experiments are conducted on five languages: Polish, Portuguese, Spanish, French, and English. HuBERT-based CAE model achieves the best results for word discrimination in all languages, despite Hu-BERT being pre-trained on English only. Also, the HuBERT-based CAE model works well in cross-lingual settings. It outperforms MFCC-based CAE models trained on the target languages when trained on one source language and tested on target languages.
Code (1)
Tasks
Self-Supervised LearningWord EmbeddingsSimilar Papers 제목 키워드 기반
A Correspondence Variational Autoencoder for Unsupervised Acoustic Word Embeddings
We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis of search, discovery, and indexing syste…
Word EmbeddingsSelf-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder
Acoustic word embeddings (AWEs) aims to map a variable-length speech segment into a fixed-dimensional representation. High-quality AWEs should be invariant to variations, such as duration, pitch and speaker. In this pape…
Word EmbeddingsLearning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech Search
We propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search. The temporal context includes the leading and trailing word sequences of a word. We assume that there exist spo…
Dynamic Time WarpingTripletWord EmbeddingsMultilingual acoustic word embedding models for processing zero-resource languages
Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech…
Transfer LearningWord EmbeddingsA comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings
Many speech processing tasks involve measuring the acoustic similarity between speech segments. Acoustic word embeddings (AWE) allow for efficient comparisons by mapping speech segments of arbitrary duration to fixed-dim…
Representation LearningWord Embeddings