paper-with-me

Papers

The NTNU System at the Interspeech 2020 Non-Native Children's Speech ASR Challenge

2020-05-18 · Tien-Hong Lo, Fu-An Chao, Shi-Yan Weng, Berlin Chen

This paper describes the NTNU ASR system participating in the Interspeech 2020 Non-Native Children's Speech ASR Challenge supported by the SIG-CHILD group of ISCA. This ASR shared task is made much more challenging due to the coexisting diversity of non-native and children speaking characteristics. In the setting of closed-track evaluation, all participants were restricted to develop their systems merely based on the speech and text corpora provided by the organizer. To work around this under-resourced issue, we built our ASR system on top of CNN-TDNNF-based acoustic models, meanwhile harnessing the synergistic power of various data augmentation strategies, including both utterance- and word-level speed perturbation and spectrogram augmentation, alongside a simple yet effective data-cleansing approach. All variants of our ASR system employed an RNN-based language model to rescore the first-pass recognition hypotheses, which was trained solely on the text dataset released by the organizer. Our system with the best configuration came out in second place, resulting in a word error rate (WER) of 17.59 %, while those of the top-performing, second runner-up and official baseline systems are 15.67%, 18.71%, 35.09%, respectively.

📄 PDF Abstract BibTeX arXiv:2005.08433

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDiversityLanguage Modelling

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Low Resource German ASR with Untranscribed Data Spoken by Non-native Children -- INTERSPEECH 2021 Shared Task SPAPL System

2021-06-18 · Jinhan Wang, Yunzheng Zhu, Ruchao Fan, Wei Chu 외

This paper describes the SPAPL system for the INTERSPEECH 2021 Challenge: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech in German. ~ 5 hours of transcribed data and ~ 60 hours of untranscri…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+3

Data augmentation using prosody and false starts to recognize non-native children's speech

2020-08-29 · Hemant Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo

This paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from child…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+3

Interspeech 2025 URGENT Speech Enhancement Challenge

2025-05-29 · Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler 외

There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embr…

DiversitySpeech Enhancement

Improving Children's Speech Recognition by Fine-tuning Self-supervised Adult Speech Representations

2022-11-14 · Renee Lu, Mostafa Shahin, Beena Ahmed

Children's speech recognition is a vital, yet largely overlooked domain when building inclusive speech technologies. The major challenge impeding progress in this domain is the lack of adequate child speech corpora; howe…

Self-Supervised Learningspeech-recognitionSpeech Recognition

The NTNU Taiwanese ASR System for Formosa Speech Recognition Challenge 2020

2021-04-09 · IJCLCLP 2021 6 · Fu-An Chao, Tien-Hong Lo, Shi-Yan Weng, Shih-Hsuan Chiu 외

This paper describes the NTNU ASR system participating in the Formosa Speech Recognition Challenge 2020 (FSR-2020) supported by the Formosa Speech in the Wild project (FSW). FSR-2020 aims at fostering the development of …

Data AugmentationSpeech Enhancementspeech-recognitionSpeech Recognition+1