paper-with-me

홈 › Papers

Corpus of 19th-century Czech Texts: Problems and Solutions

2014-05-01 · LREC 2014 5 · Karel Ku{\v{c}}era, Martin Stluka

Although the Czech language of the 19th century represents the roots of modern Czech and many features of the 20th- and 21st-century language cannot be properly understood without this historical background, the 19th-century Czech has not been thoroughly and consistently researched so far. The long-term project of a corpus of 19th-century Czech printed texts, currently in its third year, is intended to stimulate the research as well as to provide a firm material basis for it. The reason why, in our opinion, the project is worth mentioning is that it is faced with an unusual concentration of problems following mostly from the fact that the 19th century was arguably the most tumultuous period in the history of Czech, as well as from the fact that Czech is a highly inflectional language with a long history of sound changes, orthography reforms and rather discontinuous development of its vocabulary. The paper will briefly characterize the general background of the problems and present the reasoning behind the solutions that have been implemented in the ongoing project.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

LiViTo: Linguistic and Visual Features Tool for Assisted Analysis of Historic Manuscripts

2020-05-01 · LREC 2020 5 · Klaus M{\"u}ller, Aleksej Tikhonov, Rol Meyer,

We propose a mixed methods approach to the identification of scribes and authors in handwritten documents, and present LiViTo, a software tool which combines linguistic insights and computer vision techniques in order to…

Clustering

Large Language Models for the Summarization of Czech Documents: From History to the Present

2025-11-24 · Václav Tran, Jakub Šmíd, Ladislav Lenc, Jean-Pierre Salmon 외 arxiv

Text summarization is the task of automatically condensing longer texts into shorter, coherent summaries while preserving the original meaning and key information. Although this task has been extensively studied in Engli…

Text Summarization

AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts

2025-09-26 · Jiří Milička, Anna Marklová, Václav Cvrček arxiv

This article presents two corpora of English and Czech texts generated with large language models (LLMs). The motivation is to create a resource for comparing human-written texts with LLM-generated text linguistically. E…

Czech Grammar Error Correction with a Large and Diverse Corpus

2022-01-14 · Jakub Náplava, Milan Straka, Jana Straková, Alexandr Rosen

We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Er…

Grammatical Error Correction

A High-Quality Web Corpus of Czech

2012-05-01 · LREC 2012 5 · Johanka Spoustov{\'a}, Miroslav Spousta

In our paper, we present main results of the Czech grant project Internet as a Language Corpus, whose aim was to build a corpus of Czech web texts and to develop and publicly release related software tools. Our corpus ma…

ArticlesMachine TranslationPOSPOS Tagging+1