paper-with-me

Papers

Training data reduction for multilingual Spoken Language Understanding systems

2021-12-01 · ICON 2021 12 · Anmol Bansal, Anjali Shenoy, Krishna Chaitanya Pappu, Kay Rottmann, Anurag Dwarakanath

Fine-tuning self-supervised pre-trained language models such as BERT has significantly improved state-of-the-art performance on natural language processing tasks. Similar finetuning setups can also be used in commercial large scale Spoken Language Understanding (SLU) systems to perform intent classification and slot tagging on user queries. Finetuning such powerful models for use in commercial systems requires large amounts of training data and compute resources to achieve high performance. This paper is a study on the different empirical methods of identifying training data redundancies for the fine tuning paradigm. Particularly, we explore rule based and semantic techniques to reduce data in a multilingual fine tuning setting and report our results on key SLU metrics. Through our experiments, we show that we can achieve on par/better performance on fine-tuning using a reduced data set as compared to a model finetuned on the entire data set.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

intent-classificationIntent ClassificationSpoken Language Understanding

Similar Papers 제목 키워드 기반

Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition

2024-06-04 · Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi, Chin-Hui Lee

We propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and p…

Attribute

Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech

2024-06-13 · Martina Valente, Fabio Brugnara, Giovanni Morrone, Enrico Zovato 외

This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…

Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2

Code-switched inspired losses for spoken dialog representations

2021-11-01 · EMNLP 2021 11 · Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel

Spoken dialogue systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of code-switching). In this work, we introduce new pretraining losses tailored to learn gen…

RetrievalSpoken Dialogue Systems

Code-switched inspired losses for generic spoken dialog representations

2021-08-27 · Emile Chapuis, Pierre Colombo, Matthieu Labeau, Chloe Clavel

Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (\textit{e.g} in case of code-switching). In this work, we introduce new pretraining losses tailored to le…

Retrieval

Does language matter for spoken word classification? A multilingual generative meta-learning approach

2026-05-13 · Batsirayi Mupamhi Ziki, Louise Beyers, Ruan van der Merwe arxiv

Meta-learning has been shown to have better performance than supervised learning for few-shot monolingual spoken word classification. However, the meta-learning approach remains under-explored in multilingual spoken word…

Continual Learning