paper-with-me

홈 › Papers

Two Approaches to Diachronic Normalization of Polish Texts

2024-02-02 · Kacper Dudzic, Filip Graliński, Krzysztof Jassem, Marek Kubis, Piotr Wierzchoń

This paper discusses two approaches to the diachronic normalization of Polish texts: a rule-based solution that relies on a set of handcrafted patterns, and a neural normalization model based on the text-to-text transfer transformer architecture. The training and evaluation data prepared for the task are discussed in detail, along with experiments conducted to compare the proposed normalization solutions. A quantitative and qualitative analysis is made. It is shown that at the current stage of inquiry into the problem, the rule-based solution outperforms the neural one on 3 out of 4 variants of the prepared dataset, although in practice both approaches have distinct advantages and disadvantages.

📄 PDF Abstract BibTeX arXiv:2402.01300

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Using Comparable Collections of Historical Texts for Building a Diachronic Dictionary for Spelling Normalization

2013-08-01 · WS 2013 8 · Marilisa Amoia, Jose Manuel Martinez

Towards a contextualised spatial-diachronic history of literature: mapping emotional representations of the city and the country in Polish fiction from 1864 to 1939

2022-10-01 · LaTeCHCLfL (COLING) 2022 10 · Agnieszka Karlińska, Cezary Rosiński, Jan Wieczorek, Patryk Hubar 외

In this article, we discuss the conditions surrounding the building of historical and literary corpora. We describe the assumptions and method of making the original corpus of the Polish novel (1864-1939). Then, we prese…

Polish -English Statistical Machine Translation of Medical Texts

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project…

Machine TranslationPOSPOS TaggingTranslation

Geotagging a Diachronic Corpus of Alpine Texts: Comparing Distinct Approaches to Toponym Recognition

2019-09-01 · RANLP 2019 9 · Tannon Kew, Anastassia Shaitarova, Isabel Meraner, Janis Goldzycher 외

Geotagging historic and cultural texts provides valuable access to heritage data, enabling location-based searching and new geographically related discoveries. In this paper, we describe two distinct approaches to geotag…

Toponym Recognition

Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text

2023-07-21 · Lingyi Yang, Feng Jiang, Haizhou Li

The remarkable capabilities of large-scale language models, such as ChatGPT, in text generation have impressed readers and spurred researchers to devise detectors to mitigate potential risks, including misinformation, ph…

MisinformationText Generation