paper-with-me

홈 › Papers

Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation

2023-05-29 · Yui Sudo, Kazuya Hata, Kazuhiro Nakadai

End-to-end automatic speech recognition (E2E-ASR) has the potential to improve performance, but a specific issue that needs to be addressed is the difficulty it has in handling enharmonic words: named entities (NEs) with the same pronunciation and part of speech that are spelled differently. This often occurs with Japanese personal names that have the same pronunciation but different Kanji characters. Since such NE words tend to be important keywords, ASR easily loses user trust if it misrecognizes them. To solve these problems, this paper proposes a novel retraining-free customized method for E2E-ASRs based on a named-entity-aware E2E-ASR model and phoneme similarity estimation. Experimental results show that the proposed method improves the target NE character error rate by 35.7% on average relative to the conventional E2E-ASR model when selecting personal names as a target NE.

📄 PDF Abstract BibTeX arXiv:2305.17846

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Improving Neural Language Processing with Named Entities

2021-09-01 · RANLP 2021 9 · Kyoumoto Matsushita, Takuya Makino, Tomoya Iwakura

Pretraining-based neural network models have demonstrated state-of-the-art (SOTA) performances on natural language processing (NLP) tasks. The most frequently used sentence representation for neural-based NLP methods is …

Document ClassificationHeadline GenerationPart-Of-Speech TaggingPOS+2

FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition

2024-05-22 · CVPR 2024 1 · Ganggui Ding, Canyu Zhao, Wen Wang, Zhen Yang 외

Benefiting from large-scale pre-trained text-to-image (T2I) generative models, impressive progress has been achieved in customized image generation, which aims to generate user-specified concepts. Existing approaches hav…

Image Generation

An Anchor-Free Detector for Continuous Speech Keyword Spotting

2022-08-09 · Zhiyuan Zhao, Chuanxin Tang, Chengdong Yao, Chong Luo

Continuous Speech Keyword Spotting (CSKWS) is a task to detect predefined keywords in a continuous speech. In this paper, we regard CSKWS as a one-dimensional object detection task and propose a novel anchor-free detecto…

Keyword Spottingobject-detectionObject Detection

Accelerating Vision-Language Pretraining with Free Language Modeling

2023-03-24 · CVPR 2023 1 · Teng Wang, Yixiao Ge, Feng Zheng, Ran Cheng 외

The state of the arts in vision-language pretraining (VLP) achieves exemplary performance but suffers from high training costs resulting from slow convergence and long training time, especially on large-scale web dataset…

GPULanguage ModelingLanguage ModellingMasked Language Modeling+1

SIA: A Scalable Interoperable Annotation Server for Biomedical Named Entities

2020-04-08 · Johannes Kirschnick, Philippe Thomas, Roland Roller, Leonhard Hennig

Recent years showed a strong increase in biomedical sciences and an inherent increase in publication volume. Extraction of specific information from these sources requires highly sophisticated text mining and information…