paper-with-me

홈 › Papers

From Zero to Production: Baltic-Ukrainian Machine Translation Systems to Aid Refugees

2022-09-28 · Toms Bergmanis, Mārcis Pinnis

In this paper, we examine the development and usage of six low-resource machine translation systems translating between the Ukrainian language and each of the official languages of the Baltic states. We developed these systems in reaction to the escalating Ukrainian refugee crisis caused by the Russian military aggression in Ukraine in the hope that they might be helpful for refugees and public administrations. Now, two months after MT systems were made public, we analyze their usage patterns and statistics. Our findings show that the Latvian-Ukrainian and Lithuanian-Ukrainian systems are integrated into the public services of Baltic states, leading to more than 127 million translated sentences for the Lithuanian-Ukrainian system. Motivated by these findings, we further enhance our MT systems by better Ukrainian toponym translation and publish an improved version of the Lithuanian-Ukrainian system.

📄 PDF Abstract BibTeX arXiv:2209.14142

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History

2025-01-15 · Yevhen Kostiuk, Oxana Vitman, Łukasz Gagała, Artur Kiulian

In this work, we evaluated Lithuanian and general history knowledge of multilingual Large Language Models (LLMs) on a multiple-choice question-answering task. The models were tested on a dataset of Lithuanian national an…

Multiple-choiceQuestion Answering

CUNI Systems for the WMT22 Czech-Ukrainian Translation Task

2022-12-01 · Martin Popel, Jindřich Libovický, Jindřich Helcl

We present Charles University submissions to the WMT22 General Translation Shared Task on Czech-Ukrainian and Ukrainian-Czech machine translation. We present two constrained submissions based on block back-translation an…

Machine TranslationTranslation

Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages

2024-10-14 · AMTA 2022 9 · Olena Burda-Lassen

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated tex…

Machine TranslationSentenceTranslation

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty

2026-05-23 · Volodymyr Ovcharov arxiv

Tokenizer fertility the number of tokens per word imposes a hidden cost on non-English NLP. We measure fertility for ten foundation models across 25 European languages on parallel text, producing the first controlled tok…

Setting up the Data Printer with Improved English to Ukrainian Machine Translation

2024-04-23 · Yurii Paniv, Dmytro Chaplynskyi, Nikita Trynus, Volodymyr Kyrylov

To build large language models for Ukrainian we need to expand our corpora with large amounts of new algorithmic tasks expressed in natural language. Examples of task performance expressed in English are abundant, so wit…

DecoderLanguage ModelingLanguage ModellingMachine Translation+1