paper-with-me

Papers

Open-Source Conversational AI with SpeechBrain 1.0

2024-06-29 · Mirco Ravanelli, Titouan Parcollet, Adel Moumen, Sylvain de Langen, Cem Subakan, Peter Plantinga, Yingzhi Wang, Pooneh Mousavi, Luca Della Libera, Artem Ploujnikov, Francesco Paissan, Davide Borra, Salah Zaiem, Zeyu Zhao, Shucong Zhang, Georgios Karakasidis, Sung-Lin Yeh, Pierre Champion, Aku Rouhe, Rudolf Braun, Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Xuechen Liu, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaelle Laperriere, Mickael Rouvier, Renato de Mori, Yannick Esteve

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete "recipes" of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.

📄 PDF Abstract BibTeX arXiv:2407.00463

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSpeaker RecognitionSpeech Enhancementspeech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

SpeechBrain: A General-Purpose Speech Toolkit

2021-06-08 · Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe 외

SpeechBrain is an open-source and all-in-one speech toolkit. It is designed to facilitate the research and development of neural speech processing technologies by being simple, flexible, user-friendly, and well-documente…

Language IdentificationSpoken Language Understanding

CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common Voice

2023-05-29 · Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, Cem Subakan

Despite the recent advancements in Automatic Speech Recognition (ASR), the recognition of accented speech still remains a dominant problem. In order to create more inclusive ASR systems, research has shown that the integ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Classificationspeech-recognition+1

The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools

2022-06-01 · LREC 2022 6 · Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar 외

With the emergence of neural end-to-end approaches for spoken language understanding (SLU), a growing number of studies have been presented during these last three years on this topic. The major part of these works addre…

Intent DetectionSpoken Language Understanding

Timers and Such: A Practical Benchmark for Spoken Language Understanding with Numbers

2021-04-04 · Loren Lugosch, Piyush Papreja, Mirco Ravanelli, Abdelwahab Heba 외

This paper introduces Timers and Such, a new open source dataset of spoken English commands for common voice control use cases involving numbers. We describe the gap in existing spoken language understanding datasets tha…

Spoken Language Understanding

Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation

2025-10-08 · Vaibhav Srivastav, Steven Zheng, Eric Bezzam, Eustache Le Bihan 외 arxiv

We present the Open ASR Leaderboard, a reproducible benchmarking platform with community contributions from academia and industry. It compares 86 open-source and proprietary systems across 12 datasets, with English short…

Speech Recognition