paper-with-me

Papers

An Open Multilingual System for Scoring Readability of Wikipedia

2024-06-03 · Mykola Trokhymovych, Indira Sen, Martin Gerlach

With over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge. While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability of its text. However, previous investigations of the readability of Wikipedia have been restricted to English only, and there are currently no systems supporting the automatic readability assessment of the 300+ languages in Wikipedia. To bridge this gap, we develop a multilingual model to score the readability of Wikipedia articles. To train and evaluate this model, we create a novel multilingual dataset spanning 14 languages, by matching articles from Wikipedia to simplified Wikipedia and online children encyclopedias. We show that our model performs well in a zero-shot scenario, yielding a ranking accuracy of more than 80% across 14 languages and improving upon previous benchmarks. These results demonstrate the applicability of the model at scale for languages in which there is no ground-truth data available for model fine-tuning. Furthermore, we provide the first overview on the state of readability in Wikipedia beyond English.

📄 PDF Abstract BibTeX arXiv:2406.01835

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Comparing and Developing Tools to Measure the Readability of Domain-Specific Texts

2019-11-01 · IJCNLP 2019 11 · Elissa Redmiles, Lisa Maszkiewicz, Emily Hwang, Dhruv Kuchhal 외

The readability of a digital text can influence people{'}s ability to learn new things about a range topics from digital resources (e.g., Wikipedia, WebMD). Readability also impacts search rankings, and is used to evalua…

ArticlesSpecificity

Multilinguality at Your Fingertips : BabelNet, Babelfy and Beyond !

2015-06-01 · JEPTALNRECITAL 2015 6 · Roberto Navigli

Multilinguality is a key feature of today{'}s Web, and it is this feature that we leverage and exploit in our research work at the Sapienza University of Rome{'}s Linguistic Computing Laboratory, which I am going to over…

Entity LinkingSemantic SimilaritySemantic Textual SimilarityWord Sense Disambiguation

Factual Inconsistencies in Multilingual Wikipedia Tables

2025-07-24 · Silvia Cappa, Lingxiao Kong, Pille-Riin Peet, Fanfu Wei 외 arxiv

Wikipedia serves as a globally accessible knowledge source with content in over 300 languages. Despite covering the same topics, the different versions of Wikipedia are written and updated independently. This leads to fa…

Improving Multilingual Named Entity Recognition with Wikipedia Entity Type Mapping

2017-07-08 · EMNLP 2016 11 · Jian Ni, Radu Florian

The state-of-the-art named entity recognition (NER) systems are statistical machine learning models that have strong generalization capability (i.e., can recognize unseen entities that do not appear in training data) bas…

Multilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

A Topic-Aligned Multilingual Corpus of Wikipedia Articles for Studying Information Asymmetry in Low Resource Languages

2020-05-01 · LREC 2020 5 · Dwaipayan Roy, Sumit Bhatia, Prateek Jain

Wikipedia is the largest web-based open encyclopedia covering more than three hundred languages. However, different language editions of Wikipedia differ significantly in terms of their information coverage. We present a…

Articles