paper-with-me

홈 › Papers

BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge

2021-01-29 · Martin Kocour, Guillermo Cámbara, Jordi Luque, David Bonet, Mireia Farrús, Martin Karafiát, Karel Veselý, Jan ''Honza'' Ĉernocký

This paper describes joint effort of BUT and Telef\'onica Research on development of Automatic Speech Recognition systems for Albayzin 2020 Challenge. We compare approaches based on either hybrid or end-to-end models. In hybrid modelling, we explore the impact of SpecAugment layer on performance. For end-to-end modelling, we used a convolutional neural network with gated linear units (GLUs). The performance of such model is also evaluated with an additional n-gram language model to improve word error rates. We further inspect source separation methods to extract speech from noisy environment (i.e. TV shows). More precisely, we assess the effect of using a neural-based music separator named Demucs. A fusion of our best systems achieved 23.33% WER in official Albayzin 2020 evaluations. Aside from techniques used in our final submitted systems, we also describe our efforts in retrieving high quality transcripts for training.

📄 PDF Abstract BibTeX arXiv:2101.12729

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech-to-Text

Similar Papers 제목 키워드 기반

KALAKA-2: a TV Broadcast Speech Database for the Recognition of Iberian Languages in Clean and Noisy Environments

2012-05-01 · LREC 2012 5 · Luis Javier Rodr{\'\i}guez-Fuentes, Mikel Penagarikano, Amparo Varona, Mireia Diez 외

This paper presents the main features (design issues, recording setup, etc.) of KALAKA-2, a TV broadcast speech database specifically designed for the development and evaluation of language recognition systems in clean a…

KALAKA-3: a database for the recognition of spoken European languages on YouTube audios

2014-05-01 · LREC 2014 5 · Luis Javier Rodr{\'\i}guez-Fuentes, Mikel Penagarikano, Amparo Varona, Mireia Diez 외

This paper describes the main features of KALAKA-3, a speech database specifically designed for the development and evaluation of language recognition systems. The database provides TV broadcast speech for training, and …

BUT System Description to VoxCeleb Speaker Recognition Challenge 2019

2019-10-16 · Hossein Zeinali, Shuai Wang, Anna Silnova, Pavel Matějka 외

In this report, we describe the submission of Brno University of Technology (BUT) team to the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2019. We also provide a brief analysis of different systems on VoxCeleb-1 test…

Speaker Recognition

Brno Urban Dataset -- The New Data for Self-Driving Agents and Mapping Tasks

2019-09-15 · Adam Ligocki, Ales Jelinek, Ludek Zalud

Autonomous driving is a dynamically growing field of research, where quality and amount of experimental data is critical. Although several rich datasets are available these days, the demands of researchers and technical …

Autonomous Driving

Detecting Spoofing Attacks Using VGG and SincNet: BUT-Omilia Submission to ASVspoof 2019 Challenge

2019-07-13 · Hossein Zeinali, Themos Stafylakis, Georgia Athanasopoulou, Johan Rohdin 외

In this paper, we present the system description of the joint efforts of Brno University of Technology (BUT) and Omilia -- Conversational Intelligence for the ASVSpoof2019 Spoofing and Countermeasures Challenge. The prim…