paper-with-me

Papers

IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation

2021-04-16 · EMNLP 2021 11 · Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Leylia Khodra, Ayu Purwarianti, Pascale Fung

Natural language generation (NLG) benchmarks provide an important avenue to measure progress and develop better NLG systems. Unfortunately, the lack of publicly available NLG benchmarks for low-resource languages poses a challenging barrier for building NLG systems that work well for languages with limited amounts of data. Here we introduce IndoNLG, the first benchmark to measure natural language generation (NLG) progress in three low-resource -- yet widely spoken -- languages of Indonesia: Indonesian, Javanese, and Sundanese. Altogether, these languages are spoken by more than 100 million native speakers, and hence constitute an important use case of NLG systems today. Concretely, IndoNLG covers six tasks: summarization, question answering, chit-chat, and three different pairs of machine translation (MT) tasks. We collate a clean pretraining corpus of Indonesian, Sundanese, and Javanese datasets, Indo4B-Plus, which is used to pretrain our models: IndoBART and IndoGPT. We show that IndoBART and IndoGPT achieve competitive performance on all tasks -- despite using only one-fifth the parameters of a larger multilingual model, mBART-LARGE (Liu et al., 2020). This finding emphasizes the importance of pretraining on closely related, local languages to achieve more efficient learning and faster inference for very low-resource languages like Javanese and Sundanese.

📄 PDF Abstract BibTeX arXiv:2104.08200

Code (2)

SamuelCahyawijaya/indobenchmark-toolkit
indobenchmark/indobenchmark-toolkit

Tasks

Machine TranslationQuestion AnsweringText GenerationTranslation

Similar Papers 제목 키워드 기반

IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding

2020-09-11 · Asian Chapter of the Association for Computational Linguistics 2020 · Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya 외

Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available res…

BenchmarkingDiversityNatural Language UnderstandingSentence+1

NusaCrowd: Open Source Initiative for Indonesian NLP Resources

2022-12-19 · Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata 외

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought tog…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Natural Language UnderstandingSpeech Recognition

DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives

2024-11-14 · Mohammad Rifqi Farhansyah, Muhammad Zuhdi Fikri Johari, Afinzaki Amiral, Ayu Purwarianti 외

Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In t…

Optical Character RecognitionOptical Character Recognition (OCR)

NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages

2022-05-31 · Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra 외

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource langua…

Machine TranslationTranslation

NusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages

2022-07-21 · Samuel Cahyawijaya, Alham Fikri Aji, Holy Lovenia, Genta Indra Winata 외

At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarc…