Investigating Machine Learning Methods for Language and Dialect Identification of Cuneiform Texts
Identification of the languages written using cuneiform symbols is a difficult task due to the lack of resources and the problem of tokenization. The Cuneiform Language Identification task in VarDial 2019 addresses the problem of identifying seven languages and dialects written in cuneiform; Sumerian and six dialects of Akkadian language: Old Babylonian, Middle Babylonian Peripheral, Standard Babylonian, Neo-Babylonian, Late Babylonian, and Neo-Assyrian. This paper describes the approaches taken by SharifCL team to this problem in VarDial 2019. The best result belongs to an ensemble of Support Vector Machines and a naive Bayes classifier, both working on character-level features, with macro-averaged F1-score of 72.10%.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningDialect IdentificationLanguage IdentificationSimilar Papers 제목 키워드 기반
Towards spoken dialect identification of Irish
The Irish language is rich in its diversity of dialects and accents. This compounds the difficulty of creating a speech recognition system for the low-resource language, as such a system must contend with a high degree o…
Dialect IdentificationLanguage Identificationspeech-recognitionSpeech RecognitionAutomatic Arabic Dialect Identification Systems for Written Texts: A Survey
Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…
Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6Demographic Dialectal Variation in Social Media: A Case Study of African-American English
Though dialectal language is increasingly abundant on social media, few resources exist for developing NLP tools to handle such language. We conduct a case study of dialectal language in online conversational text by inv…
Dependency ParsingLanguage IdentificationComparing Pipelined and Integrated Approaches to Dialectal Arabic Neural Machine Translation
When translating diglossic languages such as Arabic, situations may arise where we would like to translate a text but do not know which dialect it is. A traditional approach to this problem is to design dialect identific…
Dialect IdentificationMachine TranslationTranslationThe Curious Case of Logistic Regression for Italian Languages and Dialects Identification
Automatic Language Identification represents an important task for improving many real-world applications such as opinion mining and machine translation. In the case of closely-related languages such as regional dialects…
Language IdentificationMachine TranslationOpinion Miningregression+1