Language Model Adaptation for Language and Dialect Identification of Text
This article describes an unsupervised language model adaptation approach that can be used to enhance the performance of language identification methods. The approach is applied to a current version of the HeLI language identification method, which is now called HeLI 2.0. We describe the HeLI 2.0 method in detail. The resulting system is evaluated using the datasets from the German dialect identification and Indo-Aryan language identification shared tasks of the VarDial workshops 2017 and 2018. The new approach with language identification provides considerably higher F1-scores than the previous HeLI method or the other systems which participated in the shared tasks. The results indicate that unsupervised language model adaptation should be considered as an option in all language identification tasks, especially in those where encountering out-of-domain data is likely.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationLanguage IdentificationLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
DADA: Dialect Adaptation via Dynamic Aggregation of Linguistic Rules
Existing large language models (LLMs) that mainly focus on Standard American English (SAE) often lead to significantly worse performance when being applied to other English dialects. While existing mitigations tackle dis…
Dialect IdentificationLanguage Discrimination and Transfer Learning for Similar Languages: Experiments with Feature Combinations and Adaptation
This paper describes the work done by team tearsofjoy participating in the VarDial 2019 Evaluation Campaign. We developed two systems based on Support Vector Machines: SVM with a flat combination of features and SVM ense…
Dialect IdentificationLanguage IdentificationTransfer LearningAutomatic Arabic Dialect Identification Systems for Written Texts: A Survey
Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…
Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6Unsupervised Deep Language and Dialect Identification for Short Texts
Automatic Language Identification (LI) or Dialect Identification (DI) of short texts of closely related languages or dialects, is one of the primary steps in many natural language processing pipelines. Language identific…
Dialect IdentificationLanguage IdentificationSentenceSentence EmbeddingsArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification
In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification mo…
Dialect IdentificationEnsemble LearningFeature EngineeringLanguage Modelling+1