Papers Speech Intent Classification
“Speech Intent Classification” 태그가 달린 논문 6편 · 필터 해제
Luganda Speech Intent Recognition for IoT Applications
The advent of Internet of Things (IoT) technology has generated massive interest in voice-controlled smart homes. While many voice-controlled smart home systems are designed to understand and support widely spoken langua…
intent-classificationIntent ClassificationIntent RecognitionSpeech Intent ClassificationLeveraging Large Language Models for Exploiting ASR Uncertainty
While large language models excel in a variety of natural language processing (NLP) tasks, to perform well on spoken language understanding (SLU) tasks, they must either rely on off-the-shelf automatic speech recognition…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+6Leveraging Pretrained ASR Encoders for Effective and Efficient End-to-End Speech Intent Classification and Slot Filling
We study speech intent classification and slot filling (SICSF) by proposing to use an encoder pretrained on speech recognition (ASR) to initialize an end-to-end (E2E) Conformer-Transformer model, which achieves the new s…
intent-classificationIntent ClassificationIntent Classification and Slot FillingSelf-Supervised Learning+5Efficient Sequence Transduction by Jointly Predicting Tokens and Durations
This paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly predicting both a token and its duration, i.…
Intent ClassificationIntent Classification and Slot FillingSlot FillingSpeech Intent Classification+1Skit-S2I: An Indian Accented Speech to Intent dataset
Conventional conversation assistants extract text transcripts from the speech signal using automatic speech recognition (ASR) and then predict intent from the transcriptions. Using end-to-end spoken language understandin…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+4mSLAM: Massively multilingual joint pre-training for speech and text
We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unlabeled speech and text in multiple langua…
cross-modal alignmentintent-classificationIntent ClassificationLanguage Modeling+4