Collecting Language Resources for the Latvian e-Government Machine Translation Platform
This paper describes corpora collection activity for building large machine translation systems for Latvian e-Government platform. We describe requirements for corpora, selection and assessment of data sources, collection of the public corpora and creation of new corpora from miscellaneous sources. Methodology, tools and assessment methods are also presented along with the results achieved, challenges faced and conclusions made. Several approaches to address the data scarceness are discussed. We summarize the volume of obtained corpora and provide quality metrics of MT systems trained on this data. Resulting MT systems for English-Latvian, Latvian English and Latvian Russian are integrated in the Latvian e-service portal and are freely available on website HUGO.LV. This paper can serve as a guidance for similar activities initiated in other countries, particularly in the context of European Language Resource Coordination action.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationMiscellaneousTranslationSimilar Papers 제목 키워드 기반
Identification of Multiword Expressions for Latvian and Lithuanian: Hybrid Approach
We discuss an experiment on automatic identification of bi-gram multi-word expressions in parallel Latvian and Lithuanian corpora. Raw corpora, lexical association measures (LAMs) and supervised machine learning (ML) are…
BIG-bench Machine LearningPOSToward a Comparable Corpus of Latvian, Russian and English Tweets
Twitter has become a rich source for linguistic data. Here, a possibility of building a trilingual Latvian-Russian-English corpus of tweets from Riga, Latvia is investigated. Such a corpus, once constructed, might be of …
Information RetrievalMachine TranslationTranslationPretraining and Benchmarking Modern Encoders for Latvian
Encoder-only transformers remain essential for practical NLP tasks. While recent advances in multilingual models have improved cross-lingual capabilities, low-resource languages such as Latvian remain underrepresented in…
T\=ezaurs.lv: the Largest Open Lexical Database for Latvian
We describe an extensive and versatile lexical resource for Latvian, an under-resourced Indo-European language, which we call Tezaurs (Latvian for {`}thesaurus{'}). It comprises a large explanatory dictionary of more tha…
Latvian National Corpora Collection – Korpuss.lv
LNCC is a diverse collection of Latvian language corpora representing both written and spoken language and is useful for both linguistic research and language modelling. The collection is intended to cover diverse Latvia…
Cultural Vocal Bursts Intensity PredictionLanguage Modelling