paper-with-me

Papers

TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus

2020-03-20 · LREC 2020 5 · Elisa Gugliotta, Marco Dinarelli

This article describes the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also known as Arabizi, is a spontaneous coding of Arabic dialects in Latin characters and arithmographs (numbers used as letters). This code-system was developed by Arabic-speaking users of social media in order to facilitate the writing in the Computer-Mediated Communication (CMC) and text messaging informal frameworks. There is variety in the realization of Arabish amongst dialects, and each Arabish code-system is under-resourced, in the same way as most of the Arabic dialects. In the last few years, the focus on Arabic dialects in the NLP field has considerably increased. Taking this into consideration, TArC will be a useful support for different types of analyses, computational and linguistic, as well as for NLP tools training. In this article we will describe preliminary work on the TArC semi-automatic construction process and some of the first analyses we developed on TArC. In addition, in order to provide a complete overview of the challenges faced during the building process, we will present the main Tunisian dialect characteristics and their encoding in Tunisian Arabish.

📄 PDF Abstract BibTeX arXiv:2003.09520

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TArC. Un corpus d'arabish tunisien

2020-06-01 · JEPTALNRECITAL 2020 6 · Elisa Gugliotta, Marco Dinarelli

TArC : Incrementally and Semi-Automatically Collecting a Tunisian arabish Corpus This article describes the collection process of the first morpho-syntactically annotated Tunisian arabish Corpus (TArC). Arabish is a spon…

TArC: Tunisian Arabish Corpus First complete release

2022-07-11 · Elisa Gugliotta, Marco Dinarelli

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent re…

LemmatizationPOSPOS TaggingTransliteration

TArC: Tunisian Arabish Corpus, First complete release

2022-06-01 · LREC 2022 6 · Elisa Gugliotta, Marco Dinarelli

In this paper we present the final result of a project focused on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the realization of two integrated and ind…

LemmatizationPOSPOS TaggingTransliteration

Automatically building a Tunisian Lexicon for Deverbal Nouns

2014-08-01 · WS 2014 8 · Ahmed Hamdi, N{\'u}ria Gala, Alexis Nasr
Speech Recognition

A Corpus and Phonetic Dictionary for Tunisian Arabic Speech Recognition

2014-05-01 · LREC 2014 5 · Abir Masmoudi, Mariem Ellouze Khmekhem, Yannick Est{\`e}ve, lamia hadrich belguith 외

In this paper we describe an effort to create a corpus and phonetic dictionary for Tunisian Arabic Automatic Speech Recognition (ASR). The corpus, named TARIC (Tunisian Arabic Railway Interaction Corpus) has a collection…

Arabic Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1