paper-with-me

Papers

ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets

2024-06-12 · Jiatong Shi, Shih-Heng Wang, William Chen, Martijn Bartelds, Vanya Bannihatti Kumar, Jinchuan Tian, Xuankai Chang, Dan Jurafsky, Karen Livescu, Hung-Yi Lee, Shinji Watanabe

ML-SUPERB evaluates self-supervised learning (SSL) models on the tasks of language identification and automatic speech recognition (ASR). This benchmark treats the models as feature extractors and uses a single shallow downstream model, which can be fine-tuned for a downstream task. However, real-world use cases may require different configurations. This paper presents ML-SUPERB~2.0, which is a new benchmark for evaluating pre-trained SSL and supervised speech models across downstream models, fine-tuning setups, and efficient model adaptation approaches. We find performance improvements over the setup of ML-SUPERB. However, performance depends on the downstream model design. Also, we find large performance differences between languages and datasets, suggesting the need for more targeted approaches to improve multilingual ASR performance.

📄 PDF Abstract BibTeX arXiv:2406.08641

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingLanguage IdentificationSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties

2025-09-08 · William Chen, Chutong Meng, Jiatong Shi, Martijn Bartelds 외 arxiv

Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. W…

ML-SUPERB: Multilingual Speech Universal PERformance Benchmark

2023-05-18 · Jiatong Shi, Dan Berrebbi, William Chen, Ho-Lam Chung 외

Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers Engl…

Automatic Speech RecognitionLanguage IdentificationSelf-Supervised Learningspeech-recognition+1

A Large-Scale Evaluation of Speech Foundation Models

2024-04-15 · Shu-wen Yang, Heng-Jui Chang, Zili Huang, Andy T. Liu 외

The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has pr…

Benchmarking

Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond

2023-10-09 · Jiatong Shi, William Chen, Dan Berrebbi, Hsiu-Hsuan Wang 외

The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identific…

Language Identificationspeech-recognitionSpeech Recognition

Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC

2025-05-30 · Qingzheng Wang, Jiancheng Sun, Yifan Peng, Shinji Watanabe

Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (A…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Identification+2