paper-with-me

Papers

GeezSwitch: Language Identification in Typologically Related Low-resourced East African Languages

2022-06-01 · LREC 2022 6 · Fitsum Gaim, Wonsuk Yang, Jong C. Park

Language identification is one of the fundamental tasks in natural language processing that is a prerequisite to data processing and numerous applications. Low-resourced languages with similar typologies are generally confused with each other in real-world applications such as machine translation, affecting the user’s experience. In this work, we present a language identification dataset for five typologically and phylogenetically related low-resourced East African languages that use the Ge’ez script as a writing system; namely Amharic, Blin, Ge’ez, Tigre, and Tigrinya. The dataset is built automatically from selected data sources, but we also performed a manual evaluation to assess its quality. Our approach to constructing the dataset is cost-effective and applicable to other low-resource languages. We integrated the dataset into an existing language-identification tool and also fine-tuned several Transformer based language models, achieving very strong results in all cases. While the task of language identification is easy for the informed person, such datasets can make a difference in real-world deployments and also serve as part of a benchmark for language understanding in the target languages. The data and models are made available at https://github.com/fgaim/geezswitch.

📄 PDF Abstract BibTeX

Code (1)

fgaim/geezswitch 공식 구현

Tasks

Language IdentificationMachine Translation

Similar Papers 제목 키워드 기반

AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages

2023-11-16 · Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak 외

Despite the recent progress on scaling multilingual machine translation (MT) to several under-resourced African languages, accurately measuring this progress remains challenging, since evaluation is often performed on n-…

Machine Translation

Delexicalized Cross-lingual Dependency Parsing for Xibe

2021-09-01 · RANLP 2021 9 · He Zhou, Sandra Kübler

Manually annotating a treebank is time-consuming and labor-intensive. We conduct delexicalized cross-lingual dependency parsing experiments, where we train the parser on one language and test on our target language. As o…

Dependency Parsing

Findings of the Shared Task on Offensive Language Identification in Tamil, Malayalam, and Kannada

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Bharathi Raja Chakravarthi, Ruba Priyadharshini, Navya Jose, Anand Kumar M 외

Detecting offensive language in social media in local languages is critical for moderating user-generated content. Thus, the field of offensive language identification in under-resourced Tamil, Malayalam and Kannada lang…

BenchmarkingLanguage Identification

Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages

2026-02-01 · Nipuna Abeykoon, Ashen Weerathunga, Pubudu Wijesinghe, Parameswari Krishnamurthy arxiv

Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when translating into typologically divergent l…

Short Text Language Identification for Under Resourced Languages

2019-11-18 · Bernardt Duvenhage

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the …

Language Identification