paper-with-me

홈 › Papers

Confidence-based Ensembles of End-to-End Speech Recognition Models

2023-06-27 · Igor Gitman, Vitaly Lavrukhin, Aleksandr Laptev, Boris Ginsburg

The number of end-to-end speech recognition models grows every year. These models are often adapted to new domains or languages resulting in a proliferation of expert systems that achieve great results on target data, while generally showing inferior performance outside of their domain of expertise. We explore combination of such experts via confidence-based ensembles: ensembles of models where only the output of the most-confident model is used. We assume that models' target data is not available except for a small validation set. We demonstrate effectiveness of our approach with two applications. First, we show that a confidence-based ensemble of 5 monolingual models outperforms a system where model selection is performed via a dedicated language identification block. Second, we demonstrate that it is possible to combine base and adapted models to achieve strong results on both original and target data. We validate all our results on multiple datasets and model architectures.

📄 PDF Abstract BibTeX arXiv:2306.15824

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationModel Selectionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Towards the Transferable Audio Adversarial Attack via Ensemble Methods

2023-04-18 · Feng Guo, Zheng Sun, Yuxuan Chen, Lei Ju

In recent years, deep learning (DL) models have achieved significant progress in many domains, such as autonomous driving, facial recognition, and speech recognition. However, the vulnerability of deep learning models to…

Adversarial AttackAutonomous Drivingspeech-recognitionSpeech Recognition+1

Late fusion ensembles for speech recognition on diverse input audio representations

2024-12-01 · Marin Jezidžić, Matej Mihelčić

We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Distilling Knowledge from Ensembles of Acoustic Models for Joint CTC-Attention End-to-End Speech Recognition

2020-05-19 · Yan Gao, Titouan Parcollet, Nicholas Lane

Knowledge distillation has been widely used to compress existing deep learning models while preserving the performance on a wide range of applications. In the specific context of Automatic Speech Recognition (ASR), disti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillationspeech-recognition+1

Multiple Confidence Gates For Joint Training Of SE And ASR

2022-04-01 · Tianrui Wang, Weibin Zhu, Yingying Gao, Junlan Feng 외

Joint training of speech enhancement model (SE) and speech recognition model (ASR) is a common solution for robust ASR in noisy environments. SE focuses on improving the auditory quality of speech, but the enhanced featu…

Robust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Confidence-Guided Error Correction for Disordered Speech Recognition

2025-09-29 · Abner Hernandez, Tomás Arias Vergara, Andreas Maier, Paula Andrea Pérez-Toro arxiv

We investigate the use of large language models (LLMs) as post-processing modules for automatic speech recognition (ASR), focusing on their ability to perform error correction for disordered speech. In particular, we pro…

Speech Recognition