paper-with-me

Papers

Cultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages

2021-09-01 · RANLP 2021 9 · Filip Markoski, Elena Markoska, Nikola Ljubešić, Eftim Zdravevski, Ljupco Kocarev

There is a shortage of high-quality corpora for South-Slavic languages. Such corpora are useful to computer scientists and researchers in social sciences and humanities alike, focusing on numerous linguistic, content analysis, and natural language processing applications. This paper presents a workflow for mining Wikipedia content and processing it into linguistically-processed corpora, applied on the Bosnian, Bulgarian, Croatian, Macedonian, Serbian, Serbo-Croatian and Slovenian Wikipedia. We make the resulting seven corpora publicly available. We showcase these corpora by comparing the content of the underlying Wikipedias, our assumption being that the content of the Wikipedias reflects broadly the interests in various topics in these Balkan nations. We perform the content comparison by using topic modelling algorithms and various distribution comparisons. The results show that all Wikipedias are topically rather similar, with all of them covering art, culture, and literature, whereas they contain differences in geography, politics, history and science.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cultural Vocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

The most controversial topics in Wikipedia: A multilingual and geographical analysis

2013-05-23 · Taha Yasseri, Anselm Spoerri, Mark Graham, János Kertész

We present, visualize and analyse the similarities and differences between the controversial topics related to "edit wars" identified in 10 different language versions of Wikipedia. After a brief review of the related wo…

Articles

Topic Discovery in Massive Text Corpora Based on Min-Hashing

2018-07-03 · Gibran Fuentes-Pineda, Ivan Vladimir Meza-Ruiz

The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to massive text corpora, the vocabulary needs …

Topic Models

A "Perspectival" Mirror of the Elephant: Investigating Language Bias on Google, ChatGPT, YouTube, and Wikipedia

2023-03-28 · Queenie Luo, Michael J. Puett, Michael D. Smith

Contrary to Google Search's mission of delivering information from "many angles so you can form your own understanding of the world," we find that Google and its most prominent returned results - Wikipedia and YouTube - …

Cultural Vocal Bursts Intensity Prediction

Deep learning for COVID-19 topic modelling via Twitter: Alpha, Delta and Omicron

2023-02-28 · Janhavi Lande, Arti Pillay, Rohitash Chandra

Topic modelling with innovative deep learning methods has gained interest for a wide range of applications that includes COVID-19. Topic modelling can provide, psychological, social and cultural insights for understandin…

Management

Mapping WordNet Domains, WordNet Topics and Wikipedia Categories to Generate Multilingual Domain Specific Resources

2014-05-01 · LREC 2014 5 · Sp Gella, ana, Carlo Strapparava, Vivi Nastase

In this paper we present the mapping between WordNet domains and WordNet topics, and the emergent Wikipedia categories. This mapping leads to a coarse alignment between WordNet and Wikipedia, useful for producing domain-…

Text CategorizationWord Sense Disambiguation