paper-with-me

홈 › Papers

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

2024-11-08 · Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in multilingual NLP. In the context of low-resource languages, however, these quality assumptions are increasingly being scrutinised. This paper critically examines the data quality of Wikipedia in a non-English setting by subjecting it to various quality filtering techniques, revealing widespread issues such as a high percentage of one-line articles and duplicate articles. We evaluate the downstream impact of quality filtering on Wikipedia and find that data quality pruning is an effective means for resource-efficient training without hurting performance, especially for low-resource languages. Moreover, we advocate for a shift in perspective from seeking a general definition of data quality towards a more language- and task-specific one. Ultimately, we aim for this study to serve as a guide to using Wikipedia for pretraining in a multilingual setting.

📄 PDF Abstract BibTeX arXiv:2411.05527

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesMultilingual NLP

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Get Your Workload in Order: Game Theoretic Prioritization of Database Auditing

2018-01-22 · Chao Yan, Bo Li, Yevgeniy Vorobeychik, Aron Laszka 외

For enhancing the privacy protections of databases, where the increasing amount of detailed personal data is stored and processed, multiple mechanisms have been developed, such as audit logging and alert triggers, which …

Heuristic Search

ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia

2019-09-11 · Aaron Halfaker, R. Stuart Geiger

Algorithmic systems---from rule-based bots to machine learning classifiers---have a long history of supporting the essential work of content moderation and other curation work in peer production projects. From counter-va…

BIG-bench Machine Learning

The Web Is Your Oyster -- Knowledge-Intensive NLP against a Very Large Web Corpus

2021-12-18 · Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko 외

In order to address increasing demands of real-world applications, the research for knowledge-intensive NLP (KI-NLP) should advance by capturing the challenges of a truly open-domain environment: web-scale knowledge, lac…

Common Sense ReasoningRetrieval

Transforming Wikipedia into an Ontology-based Information Retrieval Search Engine for Local Experts using a Third-Party Taxonomy

2015-11-04 · Gregory Grefenstette, Karima Rafes

Wikipedia is widely used for finding general information about a wide variety of topics. Its vocation is not to provide local information. For example, it provides plot, cast, and production information about a given mov…

Information RetrievalRetrieval

Fun Facts: Automatic Trivia Fact Extraction from Wikipedia

2016-12-12 · Tsurel David, Pelleg Dan, Guy Ido, Shahaf Dafna

A significant portion of web search queries directly refers to named entities. Search engines explore various ways to improve the user experience for such queries. We suggest augmenting search results with {\em trivia fa…