paper-with-me

홈 › Papers

SYN2015: Representative Corpus of Contemporary Written Czech

2016-05-01 · LREC 2016 5 · Michal K{\v{r}}en, V{\'a}clav Cvr{\v{c}}ek, Tom{\'a}{\v{s}} {\v{C}}apka, Anna {\v{C}}erm{\'a}kov{\'a}, Milena Hn{\'a}tkov{\'a}, Lucie Chlumsk{\'a}, Tom{\'a}{\v{s}} Jel{\'\i}nek, Dominika Kov{\'a}{\v{r}}{\'\i}kov{\'a}, Vladim{\'\i}r Petkevi{\v{c}}, Pavel Proch{\'a}zka, Hana Skoumalov{\'a}, Michal {\v{S}}krabal, Petr Trune{\v{c}}ek, Pavel Vond{\v{r}}i{\v{c}}ka, Adrian Jan Zasina

The paper concentrates on the design, composition and annotation of SYN2015, a new 100-million representative corpus of contemporary written Czech. SYN2015 is a sequel of the representative corpora of the SYN series that can be described as traditional (as opposed to the web-crawled corpora), featuring cleared copyright issues, well-defined composition, reliability of annotation and high-quality text processing. At the same time, SYN2015 is designed as a reflection of the variety of written Czech text production with necessary methodological and technological enhancements that include a detailed bibliographic annotation and text classification based on an updated scheme. The corpus has been produced using a completely rebuilt text processing toolchain called SynKorp. SYN2015 is lemmatized, morphologically and syntactically annotated with state-of-the-art tools. It has been published within the framework of the Czech National Corpus and it is available via the standard corpus query interface KonText at http://kontext.korpus.cz as well as a dataset in shuffled format.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

text-classificationText Classification

Similar Papers 제목 키워드 기반

Benchmark of stylistic variation in LLM-generated texts

2025-09-12 · Jiří Milička, Anna Marklová, Václav Cvrček arxiv

This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written tex…

Attributivity and Subjectivity in Contemporary Written Czech

2021-12-01 · Quasy (SyntaxFest) 2021 12 · Miroslav Kubát, Radek Čech, Xinying Chen

The SYN-series corpora of written Czech

2014-05-01 · LREC 2014 5 · Milena Hn{\'a}tkov{\'a}, Michal K{\v{r}}en, Pavel Proch{\'a}zka, Hana Skoumalov{\'a}

The paper overviews the SYN series of synchronic corpora of written Czech compiled within the framework of the Czech National Corpus project. It describes their design and processing with a focus on the annotation, i.e. …

LemmatizationMorphological Tagging

Czech Grammar Error Correction with a Large and Diverse Corpus

2022-01-14 · Jakub Náplava, Milan Straka, Jana Straková, Alexandr Rosen

We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Er…

Grammatical Error Correction

Construction and Annotation of the Jordan Comprehensive Contemporary Arabic Corpus (JCCA)

2019-08-01 · WS 2019 8 · Majdi Sawalha, Faisal Al-Shargi, Abdallah AlShdaifat, Sane Yagi 외

To compile a modern dictionary that catalogues the words in currency, and to study linguistic patterns in the contemporary language, it is necessary to have a corpus of authentic texts that reflect current usage of the l…

TAG