paper-with-me

홈 › Papers

Contextualized Automatic Speech Recognition with Dynamic Vocabulary

2024-05-22 · Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel, Yifan Peng, Shinji Watanabe

Deep biasing (DB) enhances the performance of end-to-end automatic speech recognition (E2E-ASR) models for rare words or contextual phrases using a bias list. However, most existing methods treat bias phrases as sequences of subwords in a predefined static vocabulary. This naive sequence decomposition produces unnatural token patterns, significantly lowering their occurrence probability. More advanced techniques address this problem by expanding the vocabulary with additional modules, including the external language model shallow fusion or rescoring. However, they result in increasing the workload due to the additional modules. This paper proposes a dynamic vocabulary where bias tokens can be added during inference. Each entry in a bias list is represented as a single token, unlike a sequence of existing subword tokens. This approach eliminates the need to learn subword dependencies within the bias phrases. This method is easily applied to various architectures because it only expands the embedding and output layers in common E2E-ASR architectures. Experimental results demonstrate that the proposed method improves the bias phrase WER on English and Japanese datasets by 3.1 -- 4.9 points compared with the conventional DB method.

📄 PDF Abstract BibTeX arXiv:2405.13344

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation

2025-05-29 · Zhennan Lin, Kaixun Huang, Wei Ren, Linju Yang 외

Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromis…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition

2025-05-31 · Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel, Yifan Peng 외

Contextual biasing (CB) improves automatic speech recognition for rare and unseen phrases. Recent studies have introduced dynamic vocabulary, which represents context phrases as expandable tokens in autoregressive (AR) m…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition

2025-11-14 · Yiming Rong, Yixin Zhang, Ziyi Wang, Deyang Jiang 외 arxiv

Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowle…

Speech Recognition

An efficient text augmentation approach for contextualized Mandarin speech recognition

2024-06-14 · Naijun Zheng, Xucheng Wan, Kai Liu, Ziqing Du 외

Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent limitations of speech-text data availabil…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition

2025-09-16 · Li Fu, Yu Xin, Sunlu Zeng, Lu Fan 외 arxiv

This paper presents a Pronunciation-Aware Contextualized (PAC) framework to address two key challenges in Large Language Model (LLM)-based Automatic Speech Recognition (ASR) systems: effective pronunciation modeling and …

Reinforcement LearningSpeech Recognition