paper-with-me

홈 › Papers

Pretraining Approaches for Spoken Language Recognition: TalTech Submission to the OLR 2021 Challenge

2022-05-14 · Tanel Alumäe, Kunnar Kukk

This paper investigates different pretraining approaches to spoken language identification. The paper is based on our submission to the Oriental Language Recognition 2021 Challenge. We participated in two tracks of the challenge: constrained and unconstrained language recognition. For the constrained track, we first trained a Conformer-based encoder-decoder model for multilingual automatic speech recognition (ASR), using the provided training data that had transcripts available. The shared encoder of the multilingual ASR model was then finetuned for the language identification task. For the unconstrained task, we relied on both externally available pretrained models as well as external data: the multilingual XLSR-53 wav2vec2.0 model was finetuned on the VoxLingua107 corpus for the language recognition task, and finally finetuned on the provided target language training data, augmented with CommonVoice data. Our primary metric $C_{\rm avg}$ values on the Test set are 0.0079 for the constrained task and 0.0119 for the unconstrained task which resulted in the second place in both rankings. In post-evaluation experiments, we study the amount of target language data needed for training an accurate backend model, the importance of multilingual pretraining data, and compare different models as finetuning starting points.

📄 PDF Abstract BibTeX arXiv:2205.07083

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)AvgDecoderLanguage Identificationspeech-recognitionSpeech RecognitionSpoken language identification

Similar Papers 제목 키워드 기반

Dialect Adaptation and Data Augmentation for Low-Resource ASR: TalTech Systems for the MADASR 2023 Challenge

2023-10-26 · Tanel Alumäe, Jiaming Kong, Daniil Robnikov

This paper describes Tallinn University of Technology (TalTech) systems developed for the ASRU MADASR 2023 Challenge. The challenge focuses on automatic speech recognition of dialect-rich Indian languages with limited tr…

Automatic Speech RecognitionData AugmentationDiversityspeech-recognition+1

BEA-Base: A Benchmark for ASR of Spontaneous Hungarian

2022-02-01 · P. Mihajlik, A. Balog, T. E. Gráczi, A. Kohári 외

Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets - especially for spontaneous speech - have been practically unavailable. In this paper, we introduc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

BEA-Base: A Benchmark for ASR of Spontaneous Hungarian

2022-06-01 · LREC 2022 6 · Peter Mihajlik, Andras Balog, Tekla Etelka Graczi, Anna Kohari 외

Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets – especially for spontaneous speech – have been practically unavailable. In this paper, we introduc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge

2025-06-02 · Tanel Alumäe, Artem Fedorchenko

This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Challenge. A hybrid language identification…

Language Identificationspeech-recognitionSpeech Recognition

Recent Advances in End-to-End Spoken Language Understanding

2019-09-29 · Natalia Tomashenko, Antoine Caubriere, Yannick Esteve, Antoine Laurent 외

This work investigates spoken language understanding (SLU) systems in the scenario when the semantic information is extracted directly from the speech signal by means of a single end-to-end neural network model. Two SLU …

General Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4