paper-with-me

홈 › Papers

Setting up the Data Printer with Improved English to Ukrainian Machine Translation

2024-04-23 · Yurii Paniv, Dmytro Chaplynskyi, Nikita Trynus, Volodymyr Kyrylov

To build large language models for Ukrainian we need to expand our corpora with large amounts of new algorithmic tasks expressed in natural language. Examples of task performance expressed in English are abundant, so with a high-quality translation system our community will be enabled to curate datasets faster. To aid this goal, we introduce a recipe to build a translation system using supervised finetuning of a large pretrained language model with a noisy parallel dataset of 3M pairs of Ukrainian and English sentences followed by a second phase of training using 17K examples selected by k-fold perplexity filtering on another dataset of higher quality. Our decoder-only model named Dragoman beats performance of previous state of the art encoder-decoder models on the FLORES devtest set.

📄 PDF Abstract BibTeX arXiv:2404.15196

Code (1)

lang-uk/dragoman 공식 구현 pytorch

Tasks

DecoderLanguage ModelingLanguage ModellingMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Entropy of Ukrainian

2026-04-30 · Anton Lavreniuk, Mykyta Mudryi, Markiian Chaklosh arxiv

In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the n…

EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian

2025-05-29 · Daryna Dementieva, Nikolay Babakov, Alexander Fraser

While Ukrainian NLP has seen progress in many texts processing tasks, emotion classification remains an underexplored area with no publicly available benchmark to date. In this work, we introduce EmoBench-UA, the first a…

Emotion Classification

Ukrainian Visual Word Sense Disambiguation Benchmark

2026-03-24 · Yurii Laba, Yaryna Mohytych, Ivanna Rohulia, Halyna Kyryleyza 외 arxiv

This study presents a benchmark for evaluating the Visual Word Sense Disambiguation (Visual-WSD) task in Ukrainian. The main goal of the Visual-WSD task is to identify, with minimal contextual information, the most appro…

Word Sense Disambiguation

Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages

2024-10-14 · AMTA 2022 9 · Olena Burda-Lassen

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated tex…

Machine TranslationSentenceTranslation

Spivavtor: An Instruction Tuned Ukrainian Text Editing Model

2024-04-29 · Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni

We introduce Spivavtor, a dataset, and instruction-tuned models for text editing focused on the Ukrainian language. Spivavtor is the Ukrainian-focused adaptation of the English-only CoEdIT model. Similar to CoEdIT, Spiva…

Grammatical Error CorrectionmodelText Simplification