paper-with-me

Papers

NusaWrites: Constructing High-Quality Corpora for Underrepresented and Extremely Low-Resource Languages

2023-09-19 · Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista, Emmanuel Dave, Sarah Oktavianti, Salsabil Maulana Akbar, Jhonson Lee, Nuur Shadieq, Tjeng Wawan Cenggoro, Hanung Wahyuning Linuwih, Bryan Wilie, Galih Pradipta Muridan, Genta Indra Winata, David Moeljadi, Alham Fikri Aji, Ayu Purwarianti, Pascale Fung

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled corpora for these languages through online scraping and document translation. While these methods have proven effective and cost-efficient, we have identified limitations in the resulting corpora, including a lack of lexical diversity and cultural relevance to local communities. To address this gap, we conduct a case study on Indonesian local languages. We compare the effectiveness of online scraping, human translation, and paragraph writing by native speakers in constructing datasets. Our findings demonstrate that datasets generated through paragraph writing by native speakers exhibit superior quality in terms of lexical diversity and cultural content. In addition, we present the \datasetname{} benchmark, encompassing 12 underrepresented and extremely low-resource languages spoken by millions of individuals in Indonesia. Our empirical experiment results using existing multilingual large language models conclude the need to extend these models to more underrepresented languages. We release the NusaWrites dataset at https://github.com/IndoNLP/nusa-writes.

📄 PDF Abstract BibTeX arXiv:2309.10661

Code (1)

indonlp/nusa-writes 공식 구현 pytorch

Tasks

DiversityDocument TranslationTranslation

Similar Papers 제목 키워드 기반

Cerbero-7B: A Leap Forward in Language-Specific LLMs Through Enhanced Chat Corpus Generation and Evaluation

2023-11-27 · Federico A. Galatolo, Mario G. C. A. Cimino

This study introduces a novel approach for generating high-quality, language-specific chat corpora using a self-chat mechanism. We combine a generator LLM for creating new samples and an embedder LLM to ensure diversity.…

DiversityLanguage ModellingQuestion AnsweringSentence

Data Caricatures: On the Representation of African American Language in Pretraining Corpora

2025-03-13 · Nicholas Deas, Blake Vente, Amith Ananthram, Jessica A. Grieser 외

With a combination of quantitative experiments, human judgments, and qualitative analyses, we evaluate the quantity and quality of African American Language (AAL) representation in 12 predominantly English, open-source p…

Self-Diagnosing GAN: Diagnosing Underrepresented Samples in Generative Adversarial Networks

2021-02-24 · NeurIPS 2021 12 · Jinhee Lee, HaeRi Kim, Youngkyu Hong, Hye Won Chung

Despite remarkable performance in producing realistic samples, Generative Adversarial Networks (GANs) often produce low-quality samples near low-density regions of the data manifold, e.g., samples of minor groups. Many t…

Diversity

Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus

2025-02-25 · Samy Ouzerrout

Aligned audio corpora are fundamental to NLP technologies such as ASR and speech translation, yet they remain scarce for underrepresented languages, hindering their technological integration. This paper introduces a meth…

Speech-to-Speech TranslationTranslation

JESC: Japanese-English Subtitle Corpus

2017-10-29 · LREC 2018 5 · Reid Pryzant, Yongjoo Chung, Dan Jurafsky, Denny Britz

In this paper we describe the Japanese-English Subtitle Corpus (JESC). JESC is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue. It consists of more than 3.2 millio…

Machine TranslationTranslation