Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions
This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) system pipeline. For many low-resource and endangered languages, only single-speaker recordings may be available, demanding a need for domain and speaker-invariant language ID systems. In this memo, we show that a convolutional neural network with a Self-Attentive Pooling layer shows promising results for the language identification task.
Code (0)
등록된 구현이 없습니다.
Tasks
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationSimilar Papers 제목 키워드 기반
Language ID Prediction from Speech Using Self-Attentive Pooling
This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) sy…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationLow-Resource Spoken Language Identification Using Self-Attentive Pooling and Deep 1D Time-Channel Separable Convolutions
This memo describes NTR/TSU winning submission for Low Resource ASR challenge at Dialog2021 conference, language identification track. Spoken Language Identification (LID) is an important step in a multilingual Automated…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationUtterance-level end-to-end language identification using attention-based CNN-BLSTM
In this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level,…
Language IdentificationAttentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry information in long-form audio via a recurrent …
Domain AdaptationFormLanguage IdentificationGeneralized Stock Price Prediction for Multiple Stocks Combined with News Fusion
Predicting stock prices presents challenges in financial forecasting. While traditional approaches such as ARIMA and RNNs are prevalent, recent developments in Large Language Models (LLMs) offer alternative methodologies…
Stock Price Prediction