paper-with-me

Papers

Short Text Language Identification for Under Resourced Languages

2019-11-18 · Bernardt Duvenhage

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the 11 official South African languages some of which are similar languages. The algorithm is compared to recent approaches using test sets from previous works on South African languages as well as the Discriminating between Similar Languages (DSL) shared tasks' datasets. Remaining research opportunities and pressing concerns in evaluating and comparing LID approaches are also discussed.

📄 PDF Abstract BibTeX arXiv:1911.07555

Code (1)

praekelt/feersum-lid-shared-task 공식 구현

Tasks

Language Identification

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

CoSwID, a Code Switching Identification Method Suitable for Under-Resourced Languages

2022-06-01 · SIGUL (LREC) 2022 6 · Laurent Kevers

We propose a method for identifying monolingual textual segments in multilingual documents. It requires only a minimal number of linguistic resources – word lists and monolingual corpora – and can therefore be adapted to…

Language Identification

A Comparison of Character Neural Language Model and Bootstrapping for Language Identification in Multilingual Noisy Texts

2018-06-01 · WS 2018 6 · Wafia Adouane, Simon Dobnik, Jean-Philippe Bernardy, Nasredine Semmar

This paper seeks to examine the effect of including background knowledge in the form of character pre-trained neural language model (LM), and data bootstrapping to overcome the problem of unbalanced limited resources. As…

Language IdentificationLanguage ModelingLanguage ModellingMulti-Task Learning

Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification

2023-10-01 · Spandan Dey, Premjeet Singh, Goutam Saha

Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve gener…

Language IdentificationSpoken language identification

Findings of the Shared Task on Offensive Language Identification in Tamil, Malayalam, and Kannada

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Bharathi Raja Chakravarthi, Ruba Priyadharshini, Navya Jose, Anand Kumar M 외

Detecting offensive language in social media in local languages is critical for moderating user-generated content. Thus, the field of offensive language identification in under-resourced Tamil, Malayalam and Kannada lang…

BenchmarkingLanguage Identification

DOSA: Dravidian Code-Mixed Offensive Span Identification Dataset

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Manikandan Ravikiran, Subbiah Annamalai

This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…

Language Identification