AIDA2: A Hybrid Approach for Token and Sentence Level Dialect Identification in Arabic
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationMachine TranslationSentenceTransliterationSimilar Papers 제목 키워드 기반
Development of a Guarani - Spanish Parallel Corpus
This paper presents the development of a Guarani - Spanish parallel corpus with sentence-level alignment. The Guarani sentences of the corpus use the Jopara Guarani dialect, the dialect of Guarani spoken in Paraguay, whi…
SentenceALDi: Quantifying the Arabic Level of Dialectness of Text
Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications. To h…
ArticlesDialect IdentificationSentenceExtracting Core Claims from Scientific Articles
The number of scientific articles has grown rapidly over the years and there are no signs that this growth will slow down in the near future. Because of this, it becomes increasingly difficult to keep up with the latest …
ArticlesSentenceA Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic
This paper presents a multi-dialect, multi-genre, human annotated corpus of dialectal Arabic. We collected utterances in five Arabic dialects: Levantine, Gulf, Egyptian, Iraqi and Maghrebi. We scraped newspaper websites …
Dialect IdentificationFine-Tuning BERT with Character-Level Noise for Zero-Shot Transfer to Dialects and Closely-Related Languages
In this work, we induce character-level noise in various forms when fine-tuning BERT to enable zero-shot cross-lingual transfer to unseen dialects and languages. We fine-tune BERT on three sentence-level classification t…
Cross-Lingual TransferSentenceZero-Shot Cross-Lingual Transfer