paper-with-me

홈 › Papers

Language Identification for Austronesian Languages

2022-06-09 · LREC 2022 6 · Jonathan Dunn, Wikke Nijhof

This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. Accurate language identification is an important part of developing language resources. The approach taken in this paper combines 29 Austronesian languages with 171 non-Austronesian languages to create an evaluation set drawn from eight data sources. After evaluating six approaches to language identification, we find that a classifier based on skip-gram embeddings reaches a significantly higher performance than alternate methods. We then systematically increase the number of non-Austronesian languages in the model up to a total of 800 languages to evaluate whether an increased language inventory leads to less precise predictions for the Austronesian languages of interest. This evaluation finds that there is only a minimal impact on accuracy caused by increasing the inventory of non-Austronesian languages. Further experiments adapt these language identification models for code-switching detection, achieving high accuracy across all 29 languages.

📄 PDF Abstract BibTeX arXiv:2206.04327

Code (1)

jonathandunn/pacific_codeswitch 공식 구현

Tasks

Language Identification

Similar Papers 제목 키워드 기반

Phonological Fossils: Machine Learning Detection of Non-Mainstream Vocabulary in Sulawesi Basic Lexicon

2026-03-11 · Mukhlis Amien, Go Frendi Gunawan arxiv

Basic vocabulary in many Sulawesi Austronesian languages includes forms resisting reconstruction to any proto-form with phonological patterns inconsistent with inherited roots, but whether this non-conforming vocabulary …

Number Theory Meets Linguistics: Modelling Noun Pluralisation Across 1497 Languages Using 2-adic Metrics

2022-10-08 · Gregory Baker, Diego Molla-Aliod

A simple machine learning model of pluralisation as a linear regression problem minimising a p-adic metric substantially outperforms even the most robust of Euclidean-space regressors on languages in the Indo-European, A…

regression

Parsing in the absence of related languages: Evaluating low-resource dependency parsers on Tagalog

2020-12-01 · UDW (COLING) 2020 12 · Angelina Aquino, Franz de Leon

Cross-lingual and multilingual methods have been widely suggested as options for dependency parsing of low-resource languages; however, these typically require the use of annotated data in related high-resource languages…

Dependency Parsing

Building Text-to-Speech Systems for Resource Poor Languages

2012-05-01 · LREC 2012 5 · Nur-Hana Samsudin, Mark Lee

This paper describes research on building text-to-speech synthesis systems (TTS) for resource poor languages using available resources from other languages and describes our general approach to building cross-linguistic …

ClusteringSpeech Synthesistext-to-speechText to Speech+1

A Neural Network Approach to Create Minangkabau-Indonesia Bilingual Dictionary

2022-06-01 · SIGUL (LREC) 2022 6 · Kartika Resiandi, Yohei Murakami, Arbi Haza Nasution

Indonesia has many varieties of ethnic languages, and most come from the same language family, namely Austronesian languages. Coming from that same language family, the words in Indonesian ethnic languages are very simil…

Decoder