paper-with-me

Papers

A Corpus and Phonetic Dictionary for Tunisian Arabic Speech Recognition

2014-05-01 · LREC 2014 5 · Abir Masmoudi, Mariem Ellouze Khmekhem, Yannick Est{\`e}ve, lamia hadrich belguith, Nizar Habash

In this paper we describe an effort to create a corpus and phonetic dictionary for Tunisian Arabic Automatic Speech Recognition (ASR). The corpus, named TARIC (Tunisian Arabic Railway Interaction Corpus) has a collection of audio recordings and transcriptions from dialogues in the Tunisian Railway Transport Network. The phonetic (or pronunciation) dictionary is an important ASR component that serves as an intermediary between acoustic models and language models in ASR systems. The method proposed in this paper, to automatically generate a phonetic dictionary, is rule based. For that reason, we define a set of pronunciation rules and a lexicon of exceptions. To determine the performance of our phonetic rules, we chose to evaluate our pronunciation dictionary on two types of corpora. The word error rate of word grapheme-to-phoneme mapping is around 9{\%}.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Arabic Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

2025-11-13 · Fethi Bougares, Salima Mdhaffar, Haroun Elleuch, Yannick Estève arxiv

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of …

Speech Recognition

Sentiment Analysis of Tunisian Dialects: Linguistic Ressources and Experiments

2017-04-01 · WS 2017 4 · Salima Medhaffar, Fethi Bougares, Yannick Est{\`e}ve, Lamia Hadrich-Belguith

Dialectal Arabic (DA) is significantly different from the Arabic language taught in schools and used in written communication and formal speech (broadcast news, religion, politics, etc.). There are many existing research…

Sentiment Analysis

Text and Speech-based Tunisian Arabic Sub-Dialects Identification

2020-05-01 · LREC 2020 5 · Najla Ben Abdallah, Sam{\'e}h Kchaou, Fethi Bougares

Dialect IDentification (DID) is a challenging task, and it becomes more complicated when it is about the identification of dialects that belong to the same country. Indeed, dialects of the same country are closely relate…

Dialect Identification

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

2025-04-03 · Hedi Naouara, Jean-Pierre Lorré, Jérôme Louradour

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+2

Phonetic Inventory for an Arabic Speech Corpus

2016-05-01 · LREC 2016 5 · Nawar Halabi, Mike Wald

Corpus design for speech synthesis is a well-researched topic in languages such as English compared to Modern Standard Arabic, and there is a tendency to focus on methods to automatically generate the orthographic transc…

Speech Synthesis