paper-with-me

Papers

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

2024-08-07 · Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri, Laurent Besacier

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https://github.com/hlt-mt/Speech-MASSIVE

📄 PDF Abstract BibTeX arXiv:2408.03900

Code (1)

hlt-mt/speech-massive 공식 구현 pytorch

Tasks

BenchmarkingLanguage Identificationslot-fillingSlot FillingSpoken Language Understanding

Similar Papers 제목 키워드 기반

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

2020-07-20 · Changhan Wang, Anne Wu, Juan Pino

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to f…

Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech

2022-10-27 · Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka 외

This paper proposes Virtuoso, a massively multilingual speech-text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech RecognitionSpeech Synthesis+3

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

2025-09-17 · Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Lodagala 외 arxiv

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 uniqu…

Speech Recognition

Pseudo-Labeling for Massively Multilingual Speech Recognition

2021-10-30 · Loren Lugosch, Tatiana Likhomanenko, Gabriel Synnaeve, Ronan Collobert

Semi-supervised learning through pseudo-labeling has become a staple of state-of-the-art monolingual speech recognition systems. In this work, we extend pseudo-labeling to massively multilingual speech recognition with 6…

speech-recognitionSpeech Recognition

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

2025-01-10 · Fabian David Schmidt, Ivan Vulić, Goran Glavaš, David Ifeoluwa Adelani

While recent multilingual automatic speech recognition models claim to support thousands of languages, ASR for low-resource languages remains highly unreliable due to limited bimodal speech and text training data. Better…

Automatic Speech RecognitionClassificationintent-classificationIntent Classification+8