paper-with-me

Papers

Analysis of GlobalPhone and Ethiopian Languages Speech Corpora for Multilingual ASR

2020-05-01 · LREC 2020 5 · Martha Yifiru Tachbelie, Solomon Teferra Abate, Tanja Schultz

In this paper, we present the analysis of GlobalPhone (GP) and speech corpora of Ethiopian languages (Amharic, Tigrigna, Oromo and Wolaytta). The aim of the analysis is to select speech data from GP for the development of multilingual Automatic Speech Recognition (ASR) system for the Ethiopian languages. To this end, phonetic overlaps among GP and Ethiopian languages have been analyzed. The result of our analysis shows that there is much phonetic overlap among Ethiopian languages although they are from three different language families. From GP, Turkish, Uyghur and Croatian are found to have much overlap with the Ethiopian languages. On the other hand, Korean has less phonetic overlap with the rest of the languages. Moreover, morphological complexity of the GP and Ethiopian languages, reflected by type to token ration (TTR) and out of vocabulary (OOV) rate, has been analyzed. Both metrics indicated the morphological complexity of the languages. Korean and Amharic have been identified as extremely morphologically complex compared to the other languages. Tigrigna, Russian, Turkish, Polish, etc. are also among the morphologically complex languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

GlobalPhone: Pronunciation Dictionaries in 20 Languages

2014-05-01 · LREC 2014 5 · Tanja Schultz, Tim Schlippe

This paper describes the advances in the multilingual text and speech database GlobalPhone, a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in 20 langu…

Language IdentificationLanguage ModellingSpeaker Recognitionspeech-recognition+2

Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo, and Wolaytta

2020-07-01 · WS 2020 7 · Solomon Teferra Abate, Martha Yifiru Tachbelie, Michael Melese, Hafte Abera 외

Automatic Speech Recognition (ASR) is one of the most important technologies to help people live a better life in the 21st century. However, its development requires a big speech corpus for a language. The development of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta

2020-05-01 · LREC 2020 5 · Solomon Teferra Abate, Martha Yifiru Tachbelie, Michael Melese, Hafte Abera 외

Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life. However, its development benefits from large speech corpus. The development of such a corpus is…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Parallel Corpora for bi-lingual English-Ethiopian Languages Statistical Machine Translation

2018-08-01 · COLING 2018 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting a bi…

Machine TranslationTranslation

English-Ethiopian Languages Statistical Machine Translation

2019-08-01 · WS 2019 8 · Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha 외

In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge{'}ez. The corpora are used for conducting bi-d…

Machine TranslationTranslation