paper-with-me

홈 › Papers

Simulation, Modelling and Classification of Wiki Contributors: Spotting The Good, The Bad, and The Ugly

2024-05-29 · Silvia García Méndez, Fátima Leal, Benedita Malheiro, Juan Carlos Burguillo Rial, Bruno Veloso, Adriana E. Chis, Horacio González Vélez

Data crowdsourcing is a data acquisition process where groups of voluntary contributors feed platforms with highly relevant data ranging from news, comments, and media to knowledge and classifications. It typically processes user-generated data streams to provide and refine popular services such as wikis, collaborative maps, e-commerce sites, and social networks. Nevertheless, this modus operandi raises severe concerns regarding ill-intentioned data manipulation in adversarial environments. This paper presents a simulation, modelling, and classification approach to automatically identify human and non-human (bots) as well as benign and malign contributors by using data fabrication to balance classes within experimental data sets, data stream modelling to build and update contributor profiles and, finally, autonomic data stream classification. By employing WikiVoyage - a free worldwide wiki travel guide open to contribution from the general public - as a testbed, our approach proves to significantly boost the confidence and quality of the classifier by using a class-balanced data stream, comprising both real and synthetic data. Our empirical results show that the proposed method distinguishes between benign and malign bots as well as human contributors with a classification accuracy of up to 92 %.

📄 PDF Abstract BibTeX arXiv:2405.18845

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Travel 설명 없음

Similar Papers 제목 키워드 기반

Validation sur le Web de reformulations locales: application \`a la Wikip\'edia (Assisted Rephrasing for Wikipedia Contributors through Web-based Validation) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Houda Bouamor, Aur{\'e}lien Max, Gabriel Illouz, Anne Vilnat

Low-resourced Languages and Online Knowledge Repositories: A Need-Finding Study

2024-05-26 · Hellina Hailu Nigatu, John Canny, Sarah E. Chasins

Online Knowledge Repositories (OKRs) like Wikipedia offer communities a way to share and preserve information about themselves and their ways of living. However, for communities with low-resourced languages -- including …

Articles

Modelling Uncertainty in Collaborative Document Quality Assessment

2019-11-01 · WS 2019 11 · Aili Shen, Daniel Beck, Bahar Salehi, Jianzhong Qi 외

In the context of document quality assessment, previous work has mainly focused on predicting the quality of a document relative to a putative gold standard, without paying attention to the subjectivity of this task. To …

ArticlesDecision MakingGaussian Processes

Wikibook-Bot - Automatic Generation of a Wikipedia Book

2018-12-28 · Shahar Admati, Lior Rokach, Bracha Shapira

A Wikipedia book (known as Wikibook) is a collection of Wikipedia articles on a particular theme that is organized as a book. We propose Wikibook-Bot, a machine-learning based technique for automatically generating high …

ArticlesBIG-bench Machine LearningClustering

When expertise gone missing: Uncovering the loss of prolific contributors in Wikipedia

2021-09-21 · Paramita Das, Bhanu Prakash Reddy Guda, Debajit Chakraborty, Soumya Sarkar 외

Success of planetary-scale online collaborative platforms such as Wikipedia is hinged on active and continued participation of its voluntary contributors. The phenomenal success of Wikipedia as a valued multilingual sour…

Information RetrievalRetrieval