Short Text Language Identification for Under Resourced Languages
The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the 11 official South African languages some of which are similar languages. The algorithm is compared to recent approaches using test sets from previous works on South African languages as well as the Discriminating between Similar Languages (DSL) shared tasks' datasets. Remaining research opportunities and pressing concerns in evaluating and comparing LID approaches are also discussed.
Code (1)
Tasks
Language IdentificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CoSwID, a Code Switching Identification Method Suitable for Under-Resourced Languages
We propose a method for identifying monolingual textual segments in multilingual documents. It requires only a minimal number of linguistic resources – word lists and monolingual corpora – and can therefore be adapted to…
Language IdentificationA Comparison of Character Neural Language Model and Bootstrapping for Language Identification in Multilingual Noisy Texts
This paper seeks to examine the effect of including background knowledge in the form of character pre-trained neural language model (LM), and data bootstrapping to overcome the problem of unbalanced limited resources. As…
Language IdentificationLanguage ModelingLanguage ModellingMulti-Task LearningWavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification
Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve gener…
Language IdentificationSpoken language identificationFindings of the Shared Task on Offensive Language Identification in Tamil, Malayalam, and Kannada
Detecting offensive language in social media in local languages is critical for moderating user-generated content. Thus, the field of offensive language identification in under-resourced Tamil, Malayalam and Kannada lang…
BenchmarkingLanguage IdentificationDOSA: Dravidian Code-Mixed Offensive Span Identification Dataset
This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…
Language Identification