Deciphering Related Languages
We present a method for translating texts between close language pairs. The method does not require parallel data, and it does not require the languages to be written in the same script. We show results for six language pairs: Afrikaans/Dutch, Bosnian/Serbian, Danish/Swedish, Macedonian/Bulgarian, Malaysian/Indonesian, and Polish/Belorussian. We report BLEU scores showing our method to outperform others that do not use parallel data.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingMachine TranslationSimilar Papers 제목 키워드 기반
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior
Most undeciphered lost languages exhibit two characteristics that pose significant decipherment challenges: (1) the scripts are not fully segmented into words; (2) the closest known language is not determined. We propose…
DeciphermentAn open dataset for the evolution of oracle bone characters: EVOBC
The earliest extant Chinese characters originate from oracle bone inscriptions, which are closely related to other East Asian languages. These inscriptions hold immense value for anthropology and archaeology. However, de…
DeciphermentNeural Decipherment via Minimum-Cost Flow: from Ugaritic to Linear B
In this paper we propose a novel neural approach for automatic decipherment of lost languages. To compensate for the lack of strong supervision signal, our model design is informed by patterns in language change document…
DeciphermentPhonetic and Visual Priors for Decipherment of Informal Romanization
Informal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices …
DeciphermentInductive BiasDeciphering the Underserved: Benchmarking LLM OCR for Low-Resource Scripts
This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a ben…
BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)