paper-with-me

홈 › Papers

Autocorrect for Estonian texts: final report from project EKTB25

2024-02-18 · Agnes Luhtaru, Martin Vainikko, Krista Liin, Kais Allkivi-Metsoja, Jaagup Kippar, Pille Eslon, Mark Fishel

The project was funded in 2021-2023 by the National Programme of Estonian Language Technology. Its main aim was to develop spelling and grammar correction tools for the Estonian language. The main challenge was the very small amount of available error correction data needed for such development. To mitigate this, (1) we annotated more correction data for model training and testing, (2) we tested transfer-learning, i.e. retraining machine learning models created for other tasks, so as not to depend solely on correction data, (3) we compared the developed method and model with alternatives, including large language models. We also developed automatic evaluation, which can calculate the accuracy and yield of corrections by error category, so that the effectiveness of different methods can be compared in detail. There has been a breakthrough in large language models during the project: GPT4, a commercial language model with Estonian-language support, has been created. We took into account the existence of the model when adjusting plans and in the report we present a comparison with the ability of GPT4 to improve the Estonian language text. The final results show that the approach we have developed provides better scores than GPT4 and the result is usable but not entirely reliable yet. The report also contains ideas on how GPT4 and other major language models can be implemented in the future, focusing on open-source solutions. All results of this project are open-data/open-source, with licenses that allow them to be used for purposes including commercial ones.

📄 PDF Abstract BibTeX arXiv:2402.11671

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingTransfer Learning

Similar Papers 제목 키워드 기반

MedAutoCorrect: Image-Conditioned Autocorrection in Medical Reporting

2024-12-04 · Arnold Caleb Asiimwe, Dídac Surís, Pranav Rajpurkar, Carl Vondrick

In medical reporting, the accuracy of radiological reports, whether generated by humans or machine learning algorithms, is critical. We tackle a new task in this paper: image-conditioned autocorrection of inaccuracies wi…

Estonian Wordnet: Current State and Future Prospects

2018-01-01 · GWC 2018 1 · Heili Orav, Kadri Vare, Sirli Zupping

This paper presents Estonian Wordnet (EstWN) with its latest developments. We are focusing on the time period of 2011–2017 because during this time EstWN project was supported by the National Programme for Estonian Langu…

Neural Speech Synthesis for Estonian

2020-10-06 · Liisa Rätsep, Liisi Piits, Hille Pajupuu, Indrek Hein 외

This technical report describes the results of a collaboration between the NLP research group at the University of Tartu and the Institute of Estonian Language on improving neural speech synthesis for Estonian. The repor…

SentenceSpeech Synthesistext-to-speechText to Speech

Named Entity Recognition in Estonian 19th Century Parish Court Records

2022-06-01 · LREC 2022 6 · Siim Orasmaa, Kadri Muischnek, Kristjan Poska, Anna Edela

This paper presents a new historical language resource, a corpus of Estonian Parish Court records from the years 1821-1920, annotated for named entities (NE), and reports on named entity recognition (NER) experiments usi…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Rule-based autocorrection of Piping and Instrumentation Diagrams (P&IDs) on graphs

2025-02-18 · Lukas Schulze Balhorn, Niels Seijsener, Kevin Dao, Minji Kim 외

A piping and instrumentation diagram (P&ID) is a central reference document in chemical process engineering. Currently, chemical engineers manually review P&IDs through visual inspection to find and rectify errors. Howev…

Chemical Process