paper-with-me

Papers

Sockpuppet Detection in Wikipedia: A Corpus of Real-World Deceptive Writing for Linking Identities

2013-10-24 · LREC 2014 5 · Thamar Solorio, Ragib Hasan, Mainul Mizan

This paper describes the corpus of sockpuppet cases we gathered from Wikipedia. A sockpuppet is an online user account created with a fake identity for the purpose of covering abusive behavior and/or subverting the editing regulation process. We used a semi-automated method for crawling and curating a dataset of real sockpuppet investigation cases. To the best of our knowledge, this is the first corpus available on real-world deceptive writing. We describe the process for crawling the data and some preliminary results that can be used as baseline for benchmarking research. The dataset will be released under a Creative Commons license from our project website: http://docsig.cis.uab.edu.

📄 PDF Abstract BibTeX arXiv:1310.6772

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Detecting Sockpuppetry on Wikipedia Using Meta-Learning

2025-06-12 · Luc Raszewski, Christine de Kock

Malicious sockpuppet detection on Wikipedia is critical to preserving access to reliable information on the internet and preventing the spread of disinformation. Prior machine learning approaches rely on stylistic and me…

Meta-Learning

A Case Study of Sockpuppet Detection in Wikipedia

2013-06-01 · WS 2013 6 · Thamar Solorio, Ragib Hasan, Mainul Mizan

An Army of Me: Sockpuppets in Online Discussion Communities

2017-03-21 · Srijan Kumar, Justin Cheng, Jure Leskovec, V. S. Subrahmanian

In online discussion communities, users can interact and share information and opinions on a wide variety of topics. However, some users may create multiple identities, or sockpuppets, and engage in undesired behavior by…

Detecting Sockpuppets in Deceptive Opinion Spam

2017-03-09 · Marjan Hosseinia, Arjun Mukherjee

This paper explores the problem of sockpuppet detection in deceptive opinion spam using authorship attribution and verification approaches. Two methods are explored. The first is a feature subsampling scheme that uses th…

Authorship AttributionDiversity

Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models

2025-09-27 · Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit 외 arxiv

Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is the…