Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model
Recent advances in self-supervised speech models have shown significant improvement in many downstream tasks. However, these models predominantly centered on frame-level training objectives, which can fall short in spoken language understanding tasks that require semantic comprehension. Existing works often rely on additional speech-text data as intermediate targets, which is costly in the real-world setting. To address this challenge, we propose Pseudo-Word HuBERT (PW-HuBERT), a framework that integrates pseudo word-level targets into the training process, where the targets are derived from a visually-ground speech model, notably eliminating the need for speech-text paired data. Our experimental results on four spoken language understanding (SLU) benchmarks suggest the superiority of our model in capturing semantic information.
Code (0)
등록된 구현이 없습니다.
Tasks
modelSpoken Language UnderstandingSimilar Papers 제목 키워드 기반
Wav2Seq: Pre-training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervise…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decodernamed-entity-recognition+7Unsupervised Word Segmentation Using Temporal Gradient Pseudo-Labels
Unsupervised word segmentation in audio utterances is challenging as, in speech, there is typically no gap between words. In a preliminary experiment, we show that recent deep self-supervised features are very effective …
Pseudo LabelSegmentationLanguage-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition
We improve low-resource ASR by integrating the ideas of multilingual training and self-supervised learning. Concretely, we leverage an International Phonetic Alphabet (IPA) multilingual model to create frame-level pseudo…
DiversitySelf-Supervised Learningspeech-recognitionSpeech RecognitionPseudo Label Is Better Than Human Label
State-of-the-art automatic speech recognition (ASR) systems are trained with tens of thousands of hours of labeled speech data. Human transcription is expensive and time consuming. Factors such as the quality and consist…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Labelspeech-recognition+1Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification
Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy …
Data AugmentationDenoisingPrivacy PreservingSelf-Supervised Learning+1