paper-with-me

홈 › Papers

Considerations for Multilingual Wikipedia Research

2022-04-05 · Isaac Johnson, Emily Lescak

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in the performance of language and multimodal models have led to the inclusion of many more language editions of Wikipedia in datasets and models. Building better multilingual and multimodal models requires more than just access to expanded datasets; it also requires a better understanding of what is in the data and how this content was generated. This paper seeks to provide some background to help researchers think about what differences might arise between different language editions of Wikipedia and how that might affect their models. It details three major ways in which content differences between language editions arise (local context, community and governance, and technology) and recommendations for good practices when using multilingual and multimodal data for research and modeling.

📄 PDF Abstract BibTeX arXiv:2204.02483

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mapping WordNet Domains, WordNet Topics and Wikipedia Categories to Generate Multilingual Domain Specific Resources

2014-05-01 · LREC 2014 5 · Sp Gella, ana, Carlo Strapparava, Vivi Nastase

In this paper we present the mapping between WordNet domains and WordNet topics, and the emergent Wikipedia categories. This mapping leads to a coarse alignment between WordNet and Wikipedia, useful for producing domain-…

Text CategorizationWord Sense Disambiguation

Weakly Supervised Multilingual Causality Extraction from Wikipedia

2019-11-01 · IJCNLP 2019 11 · Chikara Hashimoto

We present a method for extracting causality knowledge from Wikipedia, such as Protectionism -{\textgreater} Trade war, where the cause and effect entities correspond to Wikipedia articles. Such causality knowledge is ea…

Articles

Fair multilingual vandalism detection system for Wikipedia

2023-06-02 · Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza-Yates 외

This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced …

Feature EngineeringLanguage ModelingLanguage ModellingMasked Language Modeling

Wiki-40B: Multilingual Language Model Dataset

2020-05-01 · LREC 2020 5 · M. Guo, y, Zihang Dai, Vr 외

We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the …

Causal Language ModelingLanguage ModelingLanguage Modellingmodel

Integrating Machine-Generated Short Descriptions into the Wikipedia Android App: A Pilot Deployment of Descartes

2026-01-12 · Marija Šakota, Dmitry Brant, Cooltey Feng, Shay Nowick 외 arxiv

Short descriptions are a key part of the Wikipedia user experience, but their coverage remains uneven across languages and topics. In previous work, we introduced Descartes, a multilingual model for generating short desc…