paper-with-me

홈 › Papers

How much is Wikipedia Lagging Behind News?

2017-03-30 · Fetahu Besnik, Anand Abhijit, Anand Avishek

Wikipedia, rich in entities and events, is an invaluable resource for various knowledge harvesting, extraction and mining tasks. Numerous resources like DBpedia, YAGO and other knowledge bases are based on extracting entity and event based knowledge from it. Online news, on the other hand, is an authoritative and rich source for emerging entities, events and facts relating to existing entities. In this work, we study the creation of entities in Wikipedia with respect to news by studying how entity and event based information flows from news to Wikipedia. We analyze the lag of Wikipedia (based on the revision history of the English Wikipedia) with 20 years of \emph{The New York Times} dataset (NYT). We model and analyze the lag of entities and events, namely their first appearance in Wikipedia and in NYT, respectively. In our extensive experimental analysis, we find that almost 20\% of the external references in entity pages are news articles encoding the importance of news to Wikipedia. Second, we observe that the entity-based lag follows a normal distribution with a high standard deviation, whereas the lag for news-based events is typically very low. Finally, we find that events are responsible for creation of emergent entities with as many as 12\% of the entities mentioned in the event page are created after the creation of the event page.

📄 PDF Abstract BibTeX arXiv:1703.10345

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

TWEETQA: A Social Media Focused Question Answering Dataset

2019-07-14 · ACL 2019 7 · Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni 외

With social media becoming increasingly pop-ular on which lots of news and real-time eventsare reported, developing automated questionanswering systems is critical to the effective-ness of many applications that rely on …

ArticlesQuestion Answering

Automated News Suggestions for Populating Wikipedia Entity Pages

2017-03-30 · Besnik Fetahu, Katja Markert, Avishek Anand

Wikipedia entity pages are a valuable source of information for direct consumption and for knowledge-base construction, update and maintenance. Facts in these entity pages are typically supported by references. Recent st…

ArticlesKnowledge Base Construction

Leveraging Wikipedia article evolution for promotional tone detection

2022-05-01 · ACL 2022 5 · Christine de Kock, Andreas Vlachos

Detecting biased language is useful for a variety of applications, such as identifying hyperpartisan news sources or flagging one-sided rhetoric. In this work we introduce WikiEvolve, a dataset for document-level promoti…

Effects of algorithmic flagging on fairness: quasi-experimental evidence from Wikipedia

2020-06-04 · Nathan TeBlunthuis, Benjamin Mako Hill, Aaron Halfaker

Online community moderators often rely on social signals such as whether or not a user has an account or a profile page as clues that users may cause problems. Reliance on these clues can lead to overprofiling bias when …

Fairness

FANG-COVID: A New Large-Scale Benchmark Dataset for Fake News Detection in German

2021-11-01 · EMNLP (FEVER) 2021 11 · Justus Mattern, Yu Qiao, Elma Kerz, Daniel Wiechmann 외

As the world continues to fight the COVID-19 pandemic, it is simultaneously fighting an ‘infodemic’ – a flood of disinformation and spread of conspiracy theories leading to health threats and the division of society. To …

ArticlesFake News Detection