Learning An Invariant Speech Representation
Recognition of speech, and in particular the ability to generalize and learn from small sets of labelled examples like humans do, depends on an appropriate representation of the acoustic input. We formulate the problem of finding robust speech features for supervised learning with small sample complexity as a problem of learning representations of the signal that are maximally invariant to intraclass transformations and deformations. We propose an extension of a theory for unsupervised learning of invariant visual representations to the auditory domain and empirically evaluate its validity for voiced speech sound classification. Our version of the theory requires the memory-based, unsupervised storage of acoustic templates -- such as specific phones or words -- together with all the transformations of each that normally occur. A quasi-invariant representation for a speech segment can be obtained by projecting it to each template orbit, i.e., the set of transformed signals, and computing the associated one-dimensional empirical probability distributions. The computations can be performed by modules of filtering and pooling, and extended to hierarchical architectures. In this paper, we apply a single-layer, multicomponent representation for phonemes and demonstrate improved accuracy and decreased sample complexity for vowel classification compared to standard spectral, cepstral and perceptual features.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationSound ClassificationVowel ClassificationSimilar Papers 제목 키워드 기반
Streaming Neural Speech Codecs through Time-Invariant Representations
Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-inv…
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clust…
ClusteringRepresentation LearningTowards adversarial learning of speaker-invariant representation for speech emotion recognition
Speech emotion recognition (SER) has attracted great attention in recent years due to the high demand for emotionally intelligent speech interfaces. Deriving speaker-invariant representations for speech emotion recogniti…
ClassificationEmotion ClassificationEmotion RecognitionRepresentation Learning+1Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering
Self-supervised speech representation models have succeeded in various tasks, but improving them for content-related problems using unlabeled data is challenging. We propose speaker-invariant clustering (Spin), a novel s…
Acoustic Unit DiscoveryClusteringGPUSelf-Supervised Learning+2Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain Adaptation
Unsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefull…
Domain AdaptationRepresentation Learningspeech-recognitionSpeech Recognition