paper-with-me

Papers

Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages

2024-04-01 · Joanito Agili Lopo, Radius Tanone

In Indonesia, local languages play an integral role in the culture. However, the available Indonesian language resources still fall into the category of limited data in the Natural Language Processing (NLP) field. This is become problematic when build NLP model for these languages. To address this gap, we introduce Bhinneka Korpus, a multilingual parallel corpus featuring five Indonesian local languages. Our goal is to enhance access and utilization of these resources, extending their reach within the country. We explained in a detail the dataset collection process and associated challenges. Additionally, we experimented with translation task using the IBM Model 1 due to data constraints. The result showed that the performance of each language already shows good indications for further development. Challenges such as lexical variation, smoothing effects, and cross-linguistic variability are discussed. We intend to evaluate the corpus using advanced NLP techniques for low-resource languages, paving the way for multilingual translation models.

📄 PDF Abstract BibTeX arXiv:2404.01009

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

NusaWrites: Constructing High-Quality Corpora for Underrepresented and Extremely Low-Resource Languages

2023-09-19 · Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista 외

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled cor…

DiversityDocument TranslationTranslation

Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus

2025-02-25 · Samy Ouzerrout

Aligned audio corpora are fundamental to NLP technologies such as ASR and speech translation, yet they remain scarce for underrepresented languages, hindering their technological integration. This paper introduces a meth…

Speech-to-Speech TranslationTranslation

NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages

2022-05-31 · Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra 외

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource langua…

Machine TranslationTranslation

From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation

2024-04-14 · Artur Kiulian, Anton Polishko, Mykola Khandoga, Oryna Chubych 외

In the rapidly advancing field of AI and NLP, generative large language models (LLMs) stand at the forefront of innovation, showcasing unparalleled abilities in text understanding and generation. However, the limited rep…

BenchmarkingDiversityLanguage ModelingLanguage Modelling

Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek

2025-08-20 · Mukhammadsaid Mamasaidov, Azizullah Aral, Abror Shopulatov, Mironshoh Inomjonov arxiv

Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of s…

Machine Translation