paper-with-me

홈 › Papers

Does Corpus Quality Really Matter for Low-Resource Languages?

2022-03-15 · Mikel Artetxe, Itziar Aldabe, Rodrigo Agerri, Olatz Perez-de-Viñaspre, Aitor Soroa

The vast majority of non-English corpora are derived from automatically filtered versions of CommonCrawl. While prior work has identified major issues on the quality of these datasets (Kreutzer et al., 2021), it is not clear how this impacts downstream performance. Taking representation learning in Basque as a case study, we explore tailored crawling (manually identifying and scraping websites with high-quality content) as an alternative to filtering CommonCrawl. Our new corpus, called EusCrawl, is similar in size to the Basque portion of popular multilingual corpora like CC100 and mC4, yet it has a much higher quality according to native annotators. For instance, 66% of documents are rated as high-quality for EusCrawl, in contrast with <33% for both mC4 and CC100. Nevertheless, we obtain similar results on downstream NLU tasks regardless of the corpus used for pre-training. Our work suggests that NLU performance in low-resource languages is not primarily constrained by the quality of the data, and other factors like corpus size and domain coverage can play a more important role.

📄 PDF Abstract BibTeX arXiv:2203.08111

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora

2024-02-12 · Surangika Ranathunga, Nisansa de Silva, Menan Velayuthan, Aloka Fernando 외

We conducted a detailed analysis on the quality of web-mined corpora for two low-resource languages (making three language pairs, English-Sinhala, English-Tamil and Sinhala-Tamil). We ranked each corpus according to a si…

Machine TranslationNMTTranslation

Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision

2026-02-03 · Pritam Kadasi, Abhishek Upperwal, Mayank Singh arxiv

Instruction tuning is now the default way to train and adapt large language models, but many instruction--input--output pairs are only weakly specified: for a given input, the same output can remain plausible under sever…

Taxonomy Beats Corpus in Similarity Identification, but Does It Matter?

2015-09-01 · RANLP 2015 9 · Minh Le, Antske Fokkens

MT Quality Estimation for Computer-assisted Translation: Does it Really Help?

2015-07-01 · IJCNLP 2015 7 · Marco Turchi, Matteo Negri, Marcello Federico
Machine TranslationTranslation

Capturing the Flow of Art History

2022-12-07 · Chenxi Ji

Do we really understand how machine classifies art styles? Historically, art is perceived and interpreted by human eyes and there are always controversial discussions over how people identify and understand art. Historia…