Multilingual and Cross-Lingual Intent Detection from Spoken Data
We present a systematic study on multilingual and cross-lingual intent detection from spoken data. The study leverages a new resource put forth in this work, termed MInDS-14, a first training and evaluation resource for the intent detection task with spoken data. It covers 14 intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties. Our key results indicate that combining machine translation models with state-of-the-art multilingual sentence encoders (e.g., LaBSE) can yield strong intent detectors in the majority of target languages covered in MInDS-14, and offer comparative analyses across different axes: e.g., zero-shot versus few-shot learning, translation direction, and impact of speech recognition. We see this work as an important step towards more inclusive development and evaluation of multilingual intent detectors from spoken data, in a much wider spectrum of languages compared to prior work.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningIntent DetectionMachine TranslationSentencespeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
HIT-SCIR at MMNLU-22: Consistency Regularization for Multilingual Spoken Language Understanding
Multilingual spoken language understanding (SLU) consists of two sub-tasks, namely intent detection and slot filling. To improve the performance of these two sub-tasks, we propose to use consistency regularization based …
Data AugmentationIntent Detectionslot-fillingSlot Filling+1LaDA: Latent Dialogue Action For Zero-shot Cross-lingual Neural Network Language Modeling
Cross-lingual adaptation has proven effective in spoken language understanding (SLU) systems with limited resources. Existing methods are frequently unsatisfactory for intent detection and slot filling, particularly for …
Intent DetectionLanguage ModelingLanguage Modellingslot-filling+2Transferring Knowledge Distillation for Multilingual Social Event Detection
Recently published graph neural networks (GNNs) show promising performance at social event detection tasks. However, most studies are oriented toward monolingual data in languages with abundant training samples. This has…
Cross-Lingual Word EmbeddingsEvent DetectionKnowledge DistillationWord EmbeddingsADIMA: Abuse Detection In Multilingual Audio
Abusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perfo…
Abuse DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
While recent multilingual automatic speech recognition models claim to support thousands of languages, ASR for low-resource languages remains highly unreliable due to limited bimodal speech and text training data. Better…
Automatic Speech RecognitionClassificationintent-classificationIntent Classification+8