paper-with-me

Papers

Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer

2023-11-15 · Jin Qiu, Lu Huang, Boyu Li, Jun Zhang, Lu Lu, Zejun Ma

Deep biasing for the Transducer can improve the recognition performance of rare words or contextual entities, which is essential in practical applications, especially for streaming Automatic Speech Recognition (ASR). However, deep biasing with large-scale rare words remains challenging, as the performance drops significantly when more distractors exist and there are words with similar grapheme sequences in the bias list. In this paper, we combine the phoneme and textual information of rare words in Transducers to distinguish words with similar pronunciation or spelling. Moreover, the introduction of training with text-only data containing more rare words benefits large-scale deep biasing. The experiments on the LibriSpeech corpus demonstrate that the proposed method achieves state-of-the-art performance on rare word error rate for different scales and levels of bias lists.

📄 PDF Abstract BibTeX arXiv:2311.08966

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Phoneme-Based Contextualization for Cross-Lingual Speech Recognition in End-to-End Models

2019-06-21 · Ke Hu, Antoine Bruguier, Tara N. Sainath, Rohit Prabhavalkar 외

Contextual automatic speech recognition, i.e., biasing recognition towards a given context (e.g. user's playlists, or contacts), is challenging in end-to-end (E2E) models. Such models maintain a limited number of candida…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

PDAF: A Phonetic Debiasing Attention Framework For Speaker Verification

2024-09-09 · Massa Baali, Abdulhamid Aldoobi, Hira Dhamyal, Rita Singh 외

Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this b…

Speaker Verification

Phoneme-aware Encoding for Prefix-tree-based Contextual ASR

2023-12-15 · Hayato Futami, Emiru Tsunoo, Yosuke Kashiwagi, Hiroaki Ogawa 외

In speech recognition applications, it is important to recognize context-specific rare words, such as proper nouns. Tree-constrained Pointer Generator (TCPGen) has shown promise for this purpose, which efficiently biases…

speech-recognitionSpeech Recognition

PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation

2025-09-04 · Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda arxiv

Automatic speech recognition (ASR) systems struggle with domain-specific named entities, especially homophones. Contextual ASR improves recognition but often fails to capture fine-grained phoneme variations due to limite…

Entity DisambiguationSpeech Recognition

Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

2023-01-20 · Yinghao Aaron Li, Cong Han, Xilin Jiang, Nima Mesgarani

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are…

text-to-speechText to Speech