Feature Hashing for Language and Dialect Identification
We evaluate feature hashing for language identification (LID), a method not previously used for this task. Using a standard dataset, we first show that while feature performance is high, LID data is highly dimensional and mostly sparse ({\textgreater}99.5{\%}) as it includes large vocabularies for many languages; memory requirements grow as languages are added. Next we apply hashing using various hash sizes, demonstrating that there is no performance loss with dimensionality reductions of up to 86{\%}. We also show that using an ensemble of low-dimension hash-based classifiers further boosts performance. Feature hashing is highly useful for LID and holds great promise for future work in this area.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationDimensionality ReductionInformation RetrievalLanguage IdentificationMachine TranslationText CategorizationSimilar Papers 제목 키워드 기반
ArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification
In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification mo…
Dialect IdentificationEnsemble LearningFeature EngineeringLanguage Modelling+1Automatic Arabic Dialect Identification Systems for Written Texts: A Survey
Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…
Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6Arabic Dialect Identification for Travel and Twitter Text
This paper presents the results of the experiments done as a part of MADAR Shared Task in WANLP 2019 on Arabic Fine-Grained Dialect Identification. Dialect Identification is one of the prominent tasks in the field of Nat…
BIG-bench Machine LearningDialect IdentificationLanguage ModelingLanguage ModellingAutomatic Dialect Detection in Arabic Broadcast Speech
We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. W…
Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition+1Vowel-based Meeteilon dialect identification using a Random Forest classifier
This paper presents a vowel-based dialect identification system for Meeteilon. For this work, a vowel dataset is created by using Meeteilon Speech Corpora available at Linguistic Data Consortium for Indian Languages (LDC…
ClassificationDialect Identification