paper-with-me

홈 › Papers

Wiki Dumps to Training Corpora: South Slavic Case

2026-04-28 · Mihailo Škorić, Cosimo Palma arxiv

This paper presents a pipeline designed to transform raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divided into two major phases. The first involves extracting and cleaning text from raw dumps of Wikipedia, Wikisource, Wikibooks, Wikinews, and Wikiquote. This step requires careful handling of raw wiki markup to isolate, first of all, textual articles, and then usable natural language text within them. The second phase addresses the challenge of questionable or low-quality articles, which are often generated from databases or structured knowledge bases. These articles are characterised by repetitive patterns, generic phrasing, and minimal to no original content. To mitigate their impact, a n-gram-based filtering strategy was employed to detect high levels of textual redundancy between articles and then remove such articles from the corpora entirely. The resulting datasets aim to provide linguistically rich texts suitable for training language models or conducting comparative research across South Slavic languages. By combining systematic extraction with quality control, this work contributes to the creation of reliable, high-information corpora that reflect the authentic cultural contexts of languages. While focused on the South Slavic case in the paper, the approach is mostly language-agnostic and can be generalised to other languages.

📄 PDF Abstract BibTeX arXiv:2604.25384

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages

2021-09-01 · RANLP 2021 9 · Filip Markoski, Elena Markoska, Nikola Ljubešić, Eftim Zdravevski 외

There is a shortage of high-quality corpora for South-Slavic languages. Such corpora are useful to computer scientists and researchers in social sciences and humanities alike, focusing on numerous linguistic, content ana…

Cultural Vocal Bursts Intensity Prediction

Corpus-Based Diacritic Restoration for South Slavic Languages

2016-05-01 · LREC 2016 5 · Nikola Ljube{\v{s}}i{\'c}, Toma{\v{z}} Erjavec, Darja Fi{\v{s}}er

In computer-mediated communication, Latin-based scripts users often omit diacritics when writing. Such text is typically easily understandable to humans but very difficult for computational processing because many words …

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

2026-01-16 · Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić arxiv

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general …

CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route

2024-12-02 · Nikola Ljubešić, Taja Kuzman, Ivana Filipović Petrović, Jelena Parizoska 외

This paper introduces the CLASSLA-Express workshop series as an innovative approach to disseminating linguistic resources and infrastructure provided by the CLASSLA Knowledge Centre for South Slavic languages and the Slo…

A Parallel English - Serbian - Bulgarian - Macedonian Lexicon of Named Entities

2022-09-01 · CLIB 2022 9 · Aleksandar Petrovski

This paper describes the creation of a parallel multilingual lexicon of named entities from English to three South Slavic languages: Serbian, Bulgarian and Macedonian, with Wikipedia as a source. The basics of the propos…

Miscellaneous