paper-with-me

Papers

An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks

2024-06-20 · Varsha Suresh, Salah Aït-Mokhtar, Caroline Brun, Ioan Calapodescu

Self-supervised learning models have revolutionized the field of speech processing. However, the process of fine-tuning these models on downstream tasks requires substantial computational resources, particularly when dealing with multiple speech-processing tasks. In this paper, we explore the potential of adapter-based fine-tuning in developing a unified model capable of effectively handling multiple spoken language processing tasks. The tasks we investigate are Automatic Speech Recognition, Phoneme Recognition, Intent Classification, Slot Filling, and Spoken Emotion Recognition. We validate our approach through a series of experiments on the SUPERB benchmark, and our results indicate that adapter-based fine-tuning enables a single encoder-decoder model to perform multiple speech processing tasks with an average improvement of 18.4% across the five target tasks while staying efficient in terms of parameter updates.

📄 PDF Abstract BibTeX arXiv:2406.14747

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionDecoderEmotion Recognitionintent-classificationIntent ClassificationPhoneme RecognitionSelf-Supervised Learningslot-fillingSlot Fillingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets

2025-07-17 · Zhichao Sheng, Shilin Zhou, Chen Gong, Zhenghua Li arxiv

Spoken Language Understanding (SLU) plays a crucial role in speech-centric multimedia applications, enabling machines to comprehend spoken language in scenarios such as meetings, interviews, and customer service interact…

Spoken Language UnderstandingSpeech RecognitionSentiment Analysis

Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation

2024-05-07 · Ryan Wong, Necati Cihan Camgoz, Richard Bowden

Automatic Sign Language Translation requires the integration of both computer vision and natural language processing to effectively bridge the communication gap between sign and spoken languages. However, the deficiency …

Gloss-free Sign Language TranslationSign Language TranslationTranslation

Injecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language Understanding

2020-10-08 · Dechuan Teng, Libo Qin, Wanxiang Che, Sendong Zhao 외

In this paper, we improve Chinese spoken language understanding (SLU) by injecting word information. Previous studies on Chinese SLU do not consider the word information, failing to detect word boundaries that are benefi…

Intent DetectionSentenceslot-fillingSlot Filling+1

Leveraging Large Language Models for Exploiting ASR Uncertainty

2023-09-09 · Pranay Dighe, Yi Su, Shangshang Zheng, Yunshu Liu 외

While large language models excel in a variety of natural language processing (NLP) tasks, to perform well on spoken language understanding (SLU) tasks, they must either rely on off-the-shelf automatic speech recognition…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+6

ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation

2023-05-29 · Ambuj Mehrish, Abhinav Ramesh Kashyap, Li Yingting, Navonil Majumder 외

There are significant challenges for speaker adaptation in text-to-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address t…

Speech Synthesistext-to-speechText to Speech