paper-with-me

홈 › Papers

Entropy of Ukrainian

2026-04-30 · Anton Lavreniuk, Mykyta Mudryi, Markiian Chaklosh arxiv

In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the next character in a sentence, he was able to approximate the entropy of the English language. Several follow-up studies by other authors have since been conducted for English, and one for Hebrew. However, to date, Shannon's experiment has never been conducted for Ukrainian. In this paper, we perform this experiment for Ukrainian by recruiting 184 volunteers using social media channels. We rely on techniques used for English to approximate the entropy value of Ukrainian. The final result is an upper bound of $H_{upper}\approx1.201$ bits per character. We compare this to the performance of current Large Language Models. The methods and code used are also documented and published, along with a discussion of the main challenges encountered.

📄 PDF Abstract BibTeX arXiv:2604.27534

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Zero to Production: Baltic-Ukrainian Machine Translation Systems to Aid Refugees

2022-09-28 · Toms Bergmanis, Mārcis Pinnis

In this paper, we examine the development and usage of six low-resource machine translation systems translating between the Ukrainian language and each of the official languages of the Baltic states. We developed these s…

Machine TranslationTranslation

The Politics of Language Choice: How the Russian-Ukrainian War Influences Ukrainians' Language Use on Twitter

2023-05-04 · Daniel Racek, Brittany I. Davidson, Paul W. Thurner, Xiao Xiang Zhu 외

The use of language is innately political and often a vehicle of cultural identity as well as the basis for nation building. Here, we examine language choice and tweeting activity of Ukrainian citizens based on more than…

Attribute

Some Notes on p(e)re-Reduplication in Bulgarian and Ukrainian: A Corpus-based Study

2022-09-01 · CLIB 2022 9 · Ivan Derzhanski, Olena Siruk

We present a comparative study of p(e)re-reduplication in Bulgarian and Ukrainian, based on material from a parallel corpus of bilingual texts. We analyse all occurrences found in the corpus of close sequences and conjun…

How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?

2026-06-08 · Kateryna Karpo, Artem Chernodub arxiv

Fine-tuned Large Language Models (LLMs) dominate in Ukrainian grammatical error correction (GEC), while API-accessed LLMs remain nearly untested on minimal-edit benchmarks. We evaluate 11 commercial LLMs from four provid…

Grammatical Error Correction

Spivavtor: An Instruction Tuned Ukrainian Text Editing Model

2024-04-29 · Aman Saini, Artem Chernodub, Vipul Raheja, Vivek Kulkarni

We introduce Spivavtor, a dataset, and instruction-tuned models for text editing focused on the Ukrainian language. Spivavtor is the Ukrainian-focused adaptation of the English-only CoEdIT model. Similar to CoEdIT, Spiva…

Grammatical Error CorrectionmodelText Simplification