paper-with-me

홈 › Papers

Low-Resource Spoken Language Identification Using Self-Attentive Pooling and Deep 1D Time-Channel Separable Convolutions

2021-05-31 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR/TSU winning submission for Low Resource ASR challenge at Dialog2021 conference, language identification track. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) system pipeline. Traditionally, the ASR task requires large volumes of labeled data that are unattainable for most of the world's languages, including most of the languages of Russia. In this memo, we show that a convolutional neural network with a Self-Attentive Pooling layer shows promising results in low-resource setting for the language identification task and set up a SOTA for the Low Resource ASR challenge dataset. Additionally, we compare the structure of confusion matrices for this and significantly more diverse VoxForge dataset and state and substantiate the hypothesis that whenever the dataset is diverse enough so that the other classification factors, like gender, age etc. are well-averaged, the confusion matrix for LID system bears the language similarity measure.

📄 PDF Abstract BibTeX arXiv:2106.00052

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Similar Papers 제목 키워드 기반

Language ID Prediction from Speech Using Self-Attentive Pooling

2021-06-01 · NAACL (SIGTYP) 2021 6 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) sy…

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions

2021-04-24 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) sy…

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

A Self-Attentive Model with Gate Mechanism for Spoken Language Understanding

2018-10-01 · EMNLP 2018 10 · Changliang Li, Liang Li, Ji Qi

Spoken Language Understanding (SLU), which typically involves intent determination and slot filling, is a core component of spoken dialogue systems. Joint learning has shown to be effective in SLU given that slot tags an…

Automatic Speech Recognition (ASR)Intent DetectionLanguage ModelingLanguage Modelling+5

Spoken Language Identification System for English-Mandarin Code-Switching Child-Directed Speech

2023-06-01 · Shashi Kant Gupta, Sushant Hiray, Prashant Kukde

This work focuses on improving the Spoken Language Identification (LangId) system for a challenge that focuses on developing robust language identification systems that are reliable for non-standard, accented (Singaporea…

DecoderLanguage IdentificationSpoken language identification

SIGTYP 2021 Shared Task: Robust Spoken Language Identification

2021-06-07 · NAACL (SIGTYP) 2021 6 · Elizabeth Salesky, Badr M. Abdullah, Sabrina J. Mielke, Elena Klyachko 외

While language identification is a fundamental speech and language processing task, for many languages and language families it remains a challenging task. For many low-resource and endangered languages this is in part d…

Domain AdaptationLanguage IdentificationSpoken language identification