paper-with-me

홈 › Papers

Design of a Tigrinya Language Speech Corpus for Speech Recognition

2018-08-01 · COLING 2018 8 · Hafte Abera, Sebsibe H/mariam

In this paper, we describe the first Tigrinya Languages speech corpora designed and development for speech recognition purposes. Tigrinya, often written as Tigrigna (ትግርኛ) /tɪˈɡrinjə/ belongs to the Semitic branch of the Afro-Asiatic languages where it shows the characteristic features of a Semitic language. It is spoken by ethnic Tigray-Tigrigna people in the Horn of Africa. The paper outlines different corpus designing process analysis of related work on speech corpora creation for different languages. The authors provide also procedures that were used for the creation of Tigrinya speech recognition corpus which is the under-resourced language. One hundred and thirty speakers, native to Tigrinya language, were recorded for training and test dataset set. Each speaker read 100 texts, which consisted of syllabically rich and balanced sentences. Ten thousand sets of sentences were used to prompt sheets. These sentences contained all of the contextual syllables and phones.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Analysis of the Ethiopic Twitter Dataset for Abusive Speech in Amharic

2019-12-09 · Seid Muhie Yimam, Abinew Ali Ayele, Chris Biemann

In this paper, we present an analysis of the first Ethiopic Twitter Dataset for the Amharic language targeted for recognizing abusive speech. The dataset has been collected since 2014 that is written in Fidel script. Sin…

Tigrinya Automatic Speech recognition with Morpheme based recognition units

2020-07-01 · WS 2020 7 · Hafte Abera, sebsibe hailemariam

The Tigrinya language is agglutinative and has a large number of inflected and derived forms of words. Therefore a Tigrinya large vocabulary continuous speech recognition system often has a large number of different unit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Natural Language Processing for Tigrinya: Current State and Future Directions

2025-07-23 · Fitsum Gaim, Jong C. Park arxiv

Despite being spoken by millions of people, Tigrinya remains severely underrepresented in Natural Language Processing (NLP) research. This work presents a comprehensive survey of NLP research for Tigrinya, analyzing over…

Part-Of-Speech TaggingCross-Lingual TransferMachine TranslationSpeech Recognition

Tigrinya Number Verbalization: Rules, Algorithm, and Implementation

2026-01-06 · Fitsum Gaim, Issayas Tesfamariam arxiv

We present a systematic formalization of Tigrinya cardinal and ordinal number verbalization, addressing a gap in computational resources for the language. This work documents the canonical rules governing the expression …

Speech Synthesis

Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages

2026-03-24 · Badr M. Abdullah, Israel Abebe Azime, Atnafu Lambebo Tonja, Jesujoba O. Alabi 외 arxiv

We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, and Wolaytta. These languages belong to th…

Language IdentificationSpeech Recognition