paper-with-me

Papers

Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC

2025-05-30 · Qingzheng Wang, Jiancheng Sun, Yifan Peng, Shinji Watanabe

Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (ASR). However, these models struggle with limited resources during fine-tuning. This paper enhances multilingual LID and ASR on ML-SUPERB 2.0 by exploring multiple strategies for adapting SFMs, including frozen upstream training, partial fine-tuning, and low-rank adaptation. Furthermore, we employ data augmentation to mitigate performance gaps in few-shot settings and introduce LID Connectionist Temporal Classification (CTC) loss for regularization. Our approach achieves a 14% relative improvement in LID accuracy and a 30% relative reduction in ASR CER over the baseline on ML-SUPERB 2.0, securing second place in the Interspeech 2025 ML-SUPERB 2.0 Challenge.

📄 PDF Abstract BibTeX arXiv:2505.24200

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Identificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets

2024-06-12 · Jiatong Shi, Shih-Heng Wang, William Chen, Martijn Bartelds 외

ML-SUPERB evaluates self-supervised learning (SSL) models on the tasks of language identification and automatic speech recognition (ASR). This benchmark treats the models as feature extractors and uses a single shallow d…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingLanguage Identification+3

ML-SUPERB: Multilingual Speech Universal PERformance Benchmark

2023-05-18 · Jiatong Shi, Dan Berrebbi, William Chen, Ho-Lam Chung 외

Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers Engl…

Automatic Speech RecognitionLanguage IdentificationSelf-Supervised Learningspeech-recognition+1

Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond

2023-10-09 · Jiatong Shi, William Chen, Dan Berrebbi, Hsiu-Hsuan Wang 외

The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identific…

Language Identificationspeech-recognitionSpeech Recognition

TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge

2025-06-02 · Tanel Alumäe, Artem Fedorchenko

This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Challenge. A hybrid language identification…

Language Identificationspeech-recognitionSpeech Recognition

SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition

2024-09-16 · Ming-Hao Hsu, Hung-Yi Lee

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficien…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationIn-Context Learning+4