paper-with-me

Papers

Multilingual sentence-level bias detection in Wikipedia

2019-09-01 · RANLP 2019 9 · Aleks, Desislava rova, Fran{\c{c}}ois Lareau, Pierre Andr{\'e} M{\'e}nard

We propose a multilingual method for the extraction of biased sentences from Wikipedia, and use it to create corpora in Bulgarian, French and English. Sifting through the revision history of the articles that at some point had been considered biased and later corrected, we retrieve the last tagged and the first untagged revisions as the before/after snapshots of what was deemed a violation of Wikipedia{'}s neutral point of view policy. We extract the sentences that were removed or rewritten in that edit. The approach yields sufficient data even in the case of relatively small Wikipedias, such as the Bulgarian one, where 62k articles produced 5k biased sentences. We evaluate our method by manually annotating 520 sentences for Bulgarian and French, and 744 for English. We assess the level of noise and analyze its sources. Finally, we exploit the data with well-known classification methods to detect biased sentences. Code and datasets are hosted at https://github.com/crim-ca/wiki-bias.

📄 PDF Abstract BibTeX

Code (1)

crim-ca/wiki-bias 공식 구현

Tasks

ArticlesBias DetectionSentence

Similar Papers 제목 키워드 기반

GeBioToolkit: Automatic Extraction of Gender-Balanced Multilingual Corpus of Wikipedia Biographies

2019-12-10 · LREC 2020 5 · Marta R. Costa-jussà, Pau Li Lin, Cristina España-Bonet

We introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite thegender inequalitiespresent in Wikipedia, the t…

Sentence

A General-Purpose Multilingual Document Encoder

2023-05-11 · Onur Galoğlu, Robert Litschko, Goran Glavaš

Massively multilingual pretrained transformers (MMTs) have tremendously pushed the state of the art on multilingual NLP and cross-lingual transfer of NLP models in particular. While a large body of work leveraged MMTs to…

Cross-Lingual TransferDocument ClassificationLong-range modelingMultilingual NLP+2

Multilingual Bias Detection and Mitigation for Indian Languages

2023-12-23 · Ankita Maity, Anubhav Sharma, Rudra Dhar, Tushar Abhishek 외

Lack of diverse perspectives causes neutrality bias in Wikipedia content leading to millions of worldwide readers getting exposed by potentially inaccurate information. Hence, neutrality bias detection and mitigation is …

Bias DetectionBinary ClassificationStyle Transfer

Fair multilingual vandalism detection system for Wikipedia

2023-06-02 · Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza-Yates 외

This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced …

Feature EngineeringLanguage ModelingLanguage ModellingMasked Language Modeling

Morfessor-enriched features and multilingual training for canonical morphological segmentation

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · Aku Rouhe, Stig-Arne Grönroos, Sami Virpioja, Mathias Creutz 외

In our submission to the SIGMORPHON 2022 Shared Task on Morpheme Segmentation, we study whether an unsupervised morphological segmentation method, Morfessor, can help in a supervised setting. Previous research has shown …

Morpheme SegmentaitonSentence