Evaluating context-invariance in unsupervised speech representations
Unsupervised speech representations have taken off, with benchmarks (SUPERB, ZeroSpeech) demonstrating major progress on semi-supervised speech recognition, speech synthesis, and speech-only language modelling. Inspiration comes from the promise of ``discovering the phonemes'' of a language or a similar low-bitrate encoding. However, one of the critical properties of phoneme transcriptions is context-invariance: the phonetic context of a speech sound can have massive influence on the way it is pronounced, while the text remains stable. This is what allows tokens of the same word to have the same transcriptions -- key to language understanding. Current benchmarks do not measure context-invariance. We develop a new version of the ZeroSpeech ABX benchmark that measures context-invariance, and apply it to recent self-supervised representations. We demonstrate that the context-independence of representations is predictive of the stability of word-level representations. We suggest research concentrate on improving context-independence of self-supervised and unsupervised representations.
Code (1)
Tasks
Language Modellingspeech-recognitionSpeech RecognitionSpeech SynthesisSimilar Papers 제목 키워드 기반
Representation Learning With Hidden Unit Clustering For Low Resource Speech Applications
The representation learning of speech, without textual resources, is an area of significant interest for many low resource speech applications. In this paper, we describe an approach to self-supervised representation lea…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringRepresentation Learning+2Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
Self-supervised learning (SSL) has emerged as a promising paradigm for learning flexible speech representations from unlabeled data. By designing pretext tasks that exploit statistical regularities, SSL models can captur…
Keyword SpottingSelf-Supervised LearningSpeaker IdentificationRobust speaker recognition using unsupervised adversarial invariance
In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervi…
speaker-diarizationSpeaker DiarizationSpeaker RecognitionSpeaker VerificationUnspeech: Unsupervised Speech Context Embeddings
We introduce "Unspeech" embeddings, which are based on unsupervised learning of context feature representations for spoken language. The embeddings were trained on up to 9500 hours of crawled English speech data without …
ClusteringLearning Robust and Multilingual Speech Representations
Unsupervised speech representation learning has shown remarkable success at finding representations that correlate with phonetic structures and improve downstream speech recognition performance. However, most research ha…
Representation Learningspeech-recognitionSpeech RecognitionSpeech Representation Learning