paper-with-me

Papers

TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge

2025-06-02 · Tanel Alumäe, Artem Fedorchenko

This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Challenge. A hybrid language identification system is used, consisting of a pretrained language embedding model and a light-weight speech recognition model with a shared encoder across languages and language-specific bigram language models. For speech recognition, three models are used, where only a single model is applied for each language, depending on the training data availability and performance on held-out data. The model set consists of a finetuned version of SeamlessM4T, MMS-1B-all with custom language adapters and MMS-zeroshot. The system obtained the top overall score in the challenge.

📄 PDF Abstract BibTeX arXiv:2506.01458

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Multi-Source Evidence Fusion for Audio Question Answering

2026-03-18 · Aivo Olev, Tanel Alumäe arxiv

Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We describe TalTech's solution to the Agen…

Question Answering

Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC

2025-05-30 · Qingzheng Wang, Jiancheng Sun, Yifan Peng, Shinji Watanabe

Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (A…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Identification+2

Dialect Adaptation and Data Augmentation for Low-Resource ASR: TalTech Systems for the MADASR 2023 Challenge

2023-10-26 · Tanel Alumäe, Jiaming Kong, Daniil Robnikov

This paper describes Tallinn University of Technology (TalTech) systems developed for the ASRU MADASR 2023 Challenge. The challenge focuses on automatic speech recognition of dialect-rich Indian languages with limited tr…

Automatic Speech RecognitionData AugmentationDiversityspeech-recognition+1

The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties

2025-09-08 · William Chen, Chutong Meng, Jiatong Shi, Martijn Bartelds 외 arxiv

Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. W…

TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024

2024-07-17 · Joonas Kalda, Tanel Alumäe, Martin Lebourdais, Hervé Bredin 외

This paper describes the submissions of team TalTech-IRIT-LIS to the DISPLACE 2024 challenge. Our team participated in the speaker diarization and language diarization tracks of the challenge. In the speaker diarization …

speaker-diarizationSpeaker DiarizationSpeech Separation