An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
Self-supervised learning models have revolutionized the field of speech processing. However, the process of fine-tuning these models on downstream tasks requires substantial computational resources, particularly when dealing with multiple speech-processing tasks. In this paper, we explore the potential of adapter-based fine-tuning in developing a unified model capable of effectively handling multiple spoken language processing tasks. The tasks we investigate are Automatic Speech Recognition, Phoneme Recognition, Intent Classification, Slot Filling, and Spoken Emotion Recognition. We validate our approach through a series of experiments on the SUPERB benchmark, and our results indicate that adapter-based fine-tuning enables a single encoder-decoder model to perform multiple speech processing tasks with an average improvement of 18.4% across the five target tasks while staying efficient in terms of parameter updates.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionDecoderEmotion Recognitionintent-classificationIntent ClassificationPhoneme RecognitionSelf-Supervised Learningslot-fillingSlot Fillingspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, interviews, and customer service interact…
Spoken Language UnderstandingSpeech RecognitionSentiment AnalysisSign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation
Automatic Sign Language Translation requires the integration of both computer vision and natural language processing to effectively bridge the communication gap between sign and spoken languages. However, the deficiency …
Gloss-free Sign Language TranslationSign Language TranslationTranslationInjecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language Understanding
In this paper, we improve Chinese spoken language understanding (SLU) by injecting word information. Previous studies on Chinese SLU do not consider the word information, failing to detect word boundaries that are benefi…
Intent DetectionSentenceslot-fillingSlot Filling+1Leveraging Large Language Models for Exploiting ASR Uncertainty
While large language models excel in a variety of natural language processing (NLP) tasks, to perform well on spoken language understanding (SLU) tasks, they must either rely on off-the-shelf automatic speech recognition…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+6ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation
There are significant challenges for speaker adaptation in text-to-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address t…
Speech Synthesistext-to-speechText to Speech