paper-with-me

홈 › Papers

Fixing Errors of the Google Voice Recognizer through Phonetic Distance Metrics

2021-02-18 · Diego Campos-Sobrino, Mario Campos-Soberanis, Iván Martínez-Chin, Víctor Uc-Cetina

Speech recognition systems for the Spanish language, such as Google's, produce errors quite frequently when used in applications of a specific domain. These errors mostly occur when recognizing words new to the recognizer's language model or ad hoc to the domain. This article presents an algorithm that uses Levenshtein distance on phonemes to reduce the speech recognizer's errors. The preliminary results show that it is possible to correct the recognizer's errors significantly by using this metric and using a dictionary of specific phrases from the domain of the application. Despite being designed for particular domains, the algorithm proposed here is of general application. The phrases that must be recognized can be explicitly defined for each application, without the algorithm having to be modified. It is enough to indicate to the algorithm the set of sentences on which it must work. The algorithm's complexity is $O(tn)$, where $t$ is the number of words in the transcript to be corrected, and $n$ is the number of phrases specific to the domain.

📄 PDF Abstract BibTeX arXiv:2102.09680

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

HOC 설명 없음

Similar Papers 제목 키워드 기반

Addressing the Selection Bias in Voice Assistance: Training Voice Assistance Model in Python with Equal Data Selection

2022-12-20 · Kashav Piya, Srijal Shrestha, Cameran Frank, Estephanos Jebessa 외

In recent times, voice assistants have become a part of our day-to-day lives, allowing information retrieval by voice synthesis, voice recognition, and natural language processing. These voice assistants can be found in …

Information RetrievalRetrievalSelection bias

Listen, Attend and Spell

2015-08-05 · William Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals

We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointl…

DecoderLanguage ModelingLanguage ModellingReading Comprehension+1

Digital Speech Algorithms for Speaker De-Identification

2022-03-08 · Stefano Marinozzi, Marcos Faundez-Zanuy

The present work is based on the COST Action IC1206 for De-identification in multimedia content. It was performed to test four algorithms of voice modifications on a speech gender recognizer to find the degree of modific…

De-identification

Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling

2020-09-06 · Songxiang Liu, Yuewen Cao, Disong Wang, Xixin Wu 외

This paper proposes an any-to-many location-relative, sequence-to-sequence (seq2seq), non-parallel voice conversion approach, which utilizes text supervision during training. In this approach, we combine a bottle-neck fe…

feature selectionspeech-recognitionSpeech RecognitionVoice Conversion

Speaker Identification Experiments Under Gender De-Identification

2022-03-09 · Marcos Faundez-Zanuy, Enric Sesa-Nogueras, Stefano Marinozzi

The present work is based on the COST Action IC1206 for De-identification in multimedia content. It was performed to test four algorithms of voice modifications on a speech gender recognizer to find the degree of modific…

De-identificationSpeaker Identification