paper-with-me

Papers

Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions

2021-04-24 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) system pipeline. For many low-resource and endangered languages, only single-speaker recordings may be available, demanding a need for domain and speaker-invariant language ID systems. In this memo, we show that a convolutional neural network with a Self-Attentive Pooling layer shows promising results for the language identification task.

📄 PDF Abstract BibTeX arXiv:2104.11985

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Similar Papers 제목 키워드 기반

Language ID Prediction from Speech Using Self-Attentive Pooling

2021-06-01 · NAACL (SIGTYP) 2021 6 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) sy…

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Low-Resource Spoken Language Identification Using Self-Attentive Pooling and Deep 1D Time-Channel Separable Convolutions

2021-05-31 · Roman Bedyakin, Nikolay Mikhaylovskiy

This memo describes NTR/TSU winning submission for Low Resource ASR challenge at Dialog2021 conference, language identification track. Spoken Language Identification (LID) is an important step in a multilingual Automated…

Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Utterance-level end-to-end language identification using attention-based CNN-BLSTM

2019-02-20 · Weicheng Cai, Danwei Cai, Shen Huang, Ming Li

In this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level,…

Language Identification

Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech

2022-02-24 · Quan Wang, Yang Yu, Jason Pelecanos, Yiling Huang 외

In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry information in long-form audio via a recurrent …

Domain AdaptationFormLanguage Identification

Generalized Stock Price Prediction for Multiple Stocks Combined with News Fusion

2026-03-08 · Pei-Jun Liao, Hung-Shin Lee, Yao-Fei Cheng, Li-Wei Chen 외 arxiv

Predicting stock prices presents challenges in financial forecasting. While traditional approaches such as ARIMA and RNNs are prevalent, recent developments in Large Language Models (LLMs) offer alternative methodologies…

Stock Price Prediction