paper-with-me

Papers

Gigafida 2.0: The Reference Corpus of Written Standard Slovene

2020-05-01 · LREC 2020 5 · Simon Krek, {\v{S}}pela Arhar Holdt, Toma{\v{z}} Erjavec, Jaka {\v{C}}ibej, Andraz Repar, Polona Gantar, Nikola Ljube{\v{s}}i{\'c}, Iztok Kosem, Kaja Dobrovoljc

We describe a new version of the Gigafida reference corpus of Slovene. In addition to updating the corpus with new material and annotating it with better tools, the focus of the upgrade was also on its transformation from a general reference corpus, which contains all language variants including non-standard language, to the corpus of standard (written) Slovene. This decision could be implemented as new corpora dedicated specifically to non-standard language emerged recently. In the new version, the whole Gigafida corpus was deduplicated for the first time, which facilitates automatic extraction of data for the purposes of compilation of new lexicographic resources such as the collocations dictionary and the thesaurus of Slovene.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The goo300k corpus of historical Slovene

2012-05-01 · LREC 2012 5 · Toma{\v{z}} Erjavec

The paper presents a gold-standard reference corpus of historical Slovene containing 1,000 sampled pages from over 80 texts, which were, for the most part, written between 1750-1900. Each page of the transcription has an…

LEMMALemmatizationOptical Character Recognition (OCR)

The Slovene BNSI Broadcast News database and reference speech corpus GOS: Towards the uniform guidelines for future work

2014-05-01 · LREC 2014 5 · Andrej {\v{Z}}gank, Ana Zwitter Vitez, Darinka Verdonik

The aim of the paper is to search for common guidelines for the future development of speech databases for less resourced languages in order to make them the most useful for both main fields of their use, linguistic rese…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Quality Estimation for Synthetic Parallel Data Generation

2014-05-01 · LREC 2014 5 · Raphael Rubino, Antonio Toral, Nikola Ljube{\v{s}}i{\'c}, Gema Ram{\'\i}rez-S{\'a}nchez

This paper presents a novel approach for parallel data generation using machine translation and quality estimation. Our study focuses on pivot-based machine translation from English to Croatian through Slovene. We genera…

Machine TranslationSentenceTranslation

Corpus vs. Lexicon Supervision in Morphosyntactic Tagging: the Case of Slovene

2016-05-01 · LREC 2016 5 · Nikola Ljube{\v{s}}i{\'c}, Toma{\v{z}} Erjavec

In this paper we present a tagger developed for inflectionally rich languages for which both a training corpus and a lexicon are available. We do not constrain the tagger by the lexicon entries, allowing both for lexicon…

Neural spell-checker: Beyond words with synthetic data generation

2024-10-30 · Matej Klemen, Martin Božič, Špela Arhar Holdt, Marko Robnik-Šikonja

Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large language models, have opened new opportuniti…

Language ModelingLanguage ModellingSynthetic Data Generation