Pre-training for Spoken Language Understanding with Joint Textual and Phonetic Representation Learning
In the traditional cascading architecture for spoken language understanding (SLU), it has been observed that automatic speech recognition errors could be detrimental to the performance of natural language understanding. End-to-end (E2E) SLU models have been proposed to directly map speech input to desired semantic frame with a single model, hence mitigating ASR error propagation. Recently, pre-training technologies have been explored for these E2E models. In this paper, we propose a novel joint textual-phonetic pre-training approach for learning spoken language representations, aiming at exploring the full potentials of phonetic information to improve SLU robustness to ASR errors. We explore phoneme labels as high-level speech features, and design and compare pre-training tasks based on conditional masked language model objectives and inter-sentence relation objectives. We also investigate the efficacy of combining textual and phonetic information during fine-tuning. Experimental results on spoken language understanding benchmarks, Fluent Speech Commands and SNIPS, show that the proposed approach significantly outperforms strong baseline models and improves robustness of spoken language understanding to ASR errors.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingNatural Language UnderstandingRepresentation LearningSentencespeech-recognitionSpeech RecognitionSpoken Language UnderstandingSimilar Papers 제목 키워드 기반
Joint Learning of Word and Label Embeddings for Sequence Labelling in Spoken Language Understanding
We propose an architecture to jointly learn word and label embeddings for slot filling in spoken language understanding. The proposed approach encodes labels using a combination of word embeddings and straightforward wor…
slot-fillingSlot FillingSpoken Language UnderstandingWord EmbeddingsJoint Online Spoken Language Understanding and Language Modeling with Recurrent Neural Networks
Speaker intent detection and semantic slot filling are two critical tasks in spoken language understanding (SLU) for dialogue systems. In this paper, we describe a recurrent neural network (RNN) model that jointly perfor…
BenchmarkingIntent DetectionLanguage ModelingLanguage Modelling+3End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders
It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E) SLU models utilize sequence-to-sequence …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dialogue Act Classificationspeech-recognition+2SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions. To boost the models' performance, various pre-training methods have been pr…
Language ModelingLanguage ModellingMasked Language ModelingSpoken Language UnderstandingJoint Learning of Dialog Act Segmentation and Recognition in Spoken Dialog Using Neural Networks
Dialog act segmentation and recognition are basic natural language understanding tasks in spoken dialog systems. This paper investigates a unified architecture for these two tasks, which aims to improve the model{'}s per…
Automatic Speech Recognition (ASR)Natural Language UnderstandingSegmentationSpeech Recognition+1