paper-with-me

홈 › Papers

Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition

2023-07-17 · Shaoshi Ling, Yuxuan Hu, Shuangbei Qian, Guoli Ye, Yao Qian, Yifan Gong, Ed Lin, Michael Zeng

Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E ASR. However, integrating a pretrained language model into an E2E speech recognition model has shown limited benefits due to the mismatches between text-based LLMs and those used in E2E ASR. In this paper, we explore an alternative approach by adapting a pretrained LLMs to speech. Our experiments on fully-formatted E2E ASR transcription tasks across various domains demonstrate that our approach can effectively leverage the strengths of pretrained LLMs to produce more readable ASR transcriptions. Our model, which is based on the pretrained large language models with either an encoder-decoder or decoder-only structure, surpasses strong ASR models such as Whisper, in terms of recognition error rate, considering formats like punctuation and capitalization as well.

📄 PDF Abstract BibTeX arXiv:2307.08234

Code (1)

openai/whisper 공식 구현 pytorch

Tasks

DecoderLanguage ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

2021-04-05 · Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi 외

In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization o…

speech-recognitionSpeech RecognitionSpeech-to-Text

SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription

2025-08-07 · Raymond Grossman, Taejin Park, Kunal Dhawan, Andrew Titus 외 arxiv

We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of th…

Speech Recognition

Snow Mountain: Dataset of Audio Recordings of The Bible in Low Resource Languages

2022-06-01 · Kavitha Raju, Anjaly V, Ryan Lish, Joel Mathew

Automatic Speech Recognition (ASR) has increasing utility in the modern world. There are a many ASR models available for languages with large amounts of training data like English. However, low-resource languages are poo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Handling Numeric Expressions in Automatic Speech Recognition

2024-07-18 · Christian Huber, Alexander Waibel

This paper addresses the problem of correctly formatting numeric expressions in automatic speech recognition (ASR) transcripts. This is challenging since the expected transcript format depends on the context, e.g., 1945 …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+5

Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition

2026-03-18 · Yuxiang Mei, Delai Qiu, Shengping Liu, Jiaen Liang 외 arxiv

Speech Large Language Models (Speech-LLMs) have emerged as a powerful approach for automatic speech recognition (ASR) by aligning speech encoders with large language models. However, adapting these systems to multilingua…

parameter-efficient fine-tuningSpeech Recognition