paper-with-me

Papers

Basis Identification for Automatic Creation of Pronunciation Lexicon for Proper Names

2014-06-05 · Sunil Kumar Kopparapu, M Laxminarayana

Development of a proper names pronunciation lexicon is usually a manual effort which can not be avoided. Grapheme to phoneme (G2P) conversion modules, in literature, are usually rule based and work best for non-proper names in a particular language. Proper names are foreign to a G2P module. We follow an optimization approach to enable automatic construction of proper names pronunciation lexicon. The idea is to construct a small orthogonal set of words (basis) which can span the set of names in a given database. We propose two algorithms for the construction of this basis. The transcription lexicon of all the proper names in a database can be produced by the manual transcription of only the small set of basis words. We first construct a cost function and show that the minimization of the cost function results in a basis. We derive conditions for convergence of this cost function and validate them experimentally on a very large proper name database. Experiments show the transcription can be achieved by transcribing a set of small number of basis words. The algorithms proposed are generic and independent of language; however performance is better if the proper names have same origin, namely, same language or geographical region.

📄 PDF Abstract BibTeX arXiv:1406.1280

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automatic language identity tagging on word and sentence-level in multilingual text sources: a case-study on Luxembourgish

2014-05-01 · LREC 2014 5 · Thomas Lavergne, Gilles Adda, Martine Adda-Decker, Lori Lamel

Luxembourgish, embedded in a multilingual context on the divide between Romance and Germanic cultures, remains one of Europe{'}s under-described languages. This is due to the fact that the written production remains rela…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+4

Mlphon: A Multifunctional Grapheme-Phoneme Conversion Tool Using Finite State Transducers

2022-09-05 · IEEE Access 2022 9 · Kavya Manohar, A R jayan, Rajeev Rajan

In this article we present the design and the development of a knowledge based computational linguistic tool, Mlphon for Malayalam language. Mlphon computationally models linguistic rules using finite state transducers a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityGrapheme-to-Phoneme Conversion+10

A Generative Model of a Pronunciation Lexicon for Hindi

2017-05-06 · Pramod Pandey, Somnath Roy

Voice browser applications in Text-to- Speech (TTS) and Automatic Speech Recognition (ASR) systems crucially depend on a pronunciation lexicon. The present paper describes the model of pronunciation lexicon of Hindi deve…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Acoustic data-driven lexicon learning based on a greedy pronunciation selection framework

2017-06-12 · Xiaohui Zhang, Vimal Manohar, Daniel Povey, Sanjeev Khudanpur

Speech recognition systems for irregularly-spelled languages like English normally require hand-written pronunciations. In this paper, we describe a system for automatically obtaining pronunciations of words for which pr…

speech-recognitionSpeech Recognition

Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing

2025-01-01 · Gaofeng Cheng, Haitian Lu, Chengxu Yang, Xuyang Wang 외

Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually desig…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition