paper-with-me

Papers

IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding

2020-09-11 · Asian Chapter of the Association for Computational Linguistics 2020 · Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, Ayu Purwarianti

Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available resources. In response, we introduce the first-ever vast resource for the training, evaluating, and benchmarking on Indonesian natural language understanding (IndoNLU) tasks. IndoNLU includes twelve tasks, ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity. The datasets for the tasks lie in different domains and styles to ensure task diversity. We also provide a set of Indonesian pre-trained models (IndoBERT) trained from a large and clean Indonesian dataset Indo4B collected from publicly available sources such as social media texts, blogs, news, and websites. We release baseline models for all twelve tasks, as well as the framework for benchmark evaluation, and thus it enables everyone to benchmark their system performances.

📄 PDF Abstract BibTeX arXiv:2009.05387

Code (3)

indobenchmark/indonlu 공식 구현 pytorch
LazarusNLP/NusaBERT pytorch
assulthoni/TweetSentimentIndoBERT pytorch

Tasks

BenchmarkingDiversityNatural Language UnderstandingSentenceSentence Classification

Similar Papers 제목 키워드 기반

VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding

2024-03-23 · Phong Nguyen-Thuan Do, Son Quoc Tran, Phu Gia Hoang, Kiet Van Nguyen 외

The success of Natural Language Understanding (NLU) benchmarks in various languages, such as GLUE for English, CLUE for Chinese, KLUE for Korean, and IndoNLU for Indonesian, has facilitated the evaluation of new NLU mode…

Natural Language Understandingtext-classificationText ClassificationTransfer Learning+3

NusaCrowd: Open Source Initiative for Indonesian NLP Resources

2022-12-19 · Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata 외

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought tog…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Natural Language UnderstandingSpeech Recognition

IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation

2021-04-16 · EMNLP 2021 11 · Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio 외

Natural language generation (NLG) benchmarks provide an important avenue to measure progress and develop better NLG systems. Unfortunately, the lack of publicly available NLG benchmarks for low-resource languages poses a…

Machine TranslationQuestion AnsweringText GenerationTranslation

DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives

2024-11-14 · Mohammad Rifqi Farhansyah, Muhammad Zuhdi Fikri Johari, Afinzaki Amiral, Ayu Purwarianti 외

Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In t…

Optical Character RecognitionOptical Character Recognition (OCR)

NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages

2022-05-31 · Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra 외

Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource langua…

Machine TranslationTranslation