paper-with-me

홈 › Papers

Text-Utilization for Encoder-dominated Speech Recognition Models

2026-04-29 · Albert Zeyer, Tim Posielek, Ralf Schlüter, Hermann Ney arxiv

This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We provide a comprehensive comparison of techniques to integrate text-only data, including modality matching and dynamic downsampling to reach text-level representations within the encoder. Our experiments on the LibriSpeech corpus show that a larger encoder with a smaller decoder can equal or surpass the performance of architectures with larger decoders. We demonstrate that simple configurations, such as random duration models, are often more effective than complex alternatives, significantly simplifying the training pipeline. All code and recipes are made publicly available.

📄 PDF Abstract BibTeX arXiv:2604.26514

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words

2024-08-15 · Kento Nozawa, Takashi Masuko, Toru Taniguchi

We develop a large language model (LLM) based automatic speech recognition (ASR) system that can be contextualized by providing keywords as prior information in text prompts. We adopt decoder-only architecture and use ou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Modeling+4

End-to-End Code-Switching ASR for Low-Resourced Language Pairs

2019-09-27 · Xianghu Yue, Grandee Lee, Emre Yilmaz, Fang Deng 외

Despite the significant progress in end-to-end (E2E) automatic speech recognition (ASR), E2E ASR for low resourced code-switching (CS) speech has not been well studied. In this work, we describe an E2E ASR pipeline for t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Transsion TSUP's speech recognition system for ASRU 2023 MADASR Challenge

2023-07-20 · Xiaoxiao Li, Gaosheng Zhang, An Zhu, Weiyong Li 외

This paper presents a speech recognition system developed by the Transsion Speech Understanding Processing Team (TSUP) for the ASRU 2023 MADASR Challenge. The system focuses on adapting ASR models for low-resource Indian…

DecoderLanguage ModelingLanguage Modellingspeech-recognition+1

Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning

2024-12-30 · Zixiang Wan, Ziyue Qiu, Yiyang Liu, Wei-Qiang Zhang

Speech Emotion Recognition (SER) involves analyzing vocal expressions to determine the emotional state of speakers, where the comprehensive and thorough utilization of audio information is paramount. Therefore, we propos…

Emotion RecognitionMulti-Task LearningSelf-Supervised LearningSpeech Emotion Recognition

Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities

2024-10-11 · Aulia Adila, Dessi Lestari, Ayu Purwarianti, Dipta Tanaya 외

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition