paper-with-me

홈 › Papers

Orthographic and Morphological Correspondences between Related Slavic Languages as a Base for Modeling of Mutual Intelligibility

2016-05-01 · LREC 2016 5 · Andrea Fischer, Kl{\'a}ra J{\'a}grov{\'a}, Irina Stenger, Tania Avgustinova, Dietrich Klakow, Rol Marti,

In an intercomprehension scenario, typically a native speaker of language L1 is confronted with output from an unknown, but related language L2. In this setting, the degree to which the receiver recognizes the unfamiliar words greatly determines communicative success. Despite exhibiting great string-level differences, cognates may be recognized very successfully if the receiver is aware of regular correspondences which allow to transform the unknown word into its familiar form. Modeling L1-L2 intercomprehension then requires the identification of all the regular correspondences between languages L1 and L2. We here present a set of linguistic orthographic correspondences manually compiled from comparative linguistics literature along with a set of statistically-inferred suggestions for correspondence rules. In order to do statistical inference, we followed the Minimum Description Length principle, which proposes to choose those rules which are most effective at describing the data. Our statistical model was able to reproduce most of our linguistic correspondences (88.5{\%} for Czech-Polish and 75.7{\%} for Bulgarian-Russian) and furthermore allowed to easily identify many more non-trivial correspondences which also cover aspects of morphology.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension

2022-06-01 · LREC 2022 6 · Irina Stenger, Philip Georgis, Tania Avgustinova, Bernd Möbius 외

We focus on the syntactic variation and measure syntactic distances between nine Slavic languages (Belarusian, Bulgarian, Croatian, Czech, Polish, Slovak, Slovene, Russian, and Ukrainian) using symmetric measures of inse…

Cloze TestReading Comprehension

Character Alignment in Morphologically Complex Translation Sets for Related Languages

2020-12-01 · VarDial (COLING) 2020 12 · Michael Gasser, Binyam Ephrem Seyoum, Nazareth Amlesom Kifle

For languages with complex morphology, word-to-word translation is a task with various potential applications, for example, in information retrieval, language instruction, and dictionary creation, as well as in machine t…

Information RetrievalMachine TranslationRetrievalTranslation+1

Exploiting Cross-Dialectal Gold Syntax for Low-Resource Historical Languages: Towards a Generic Parser for Pre-Modern Slavic

2020-11-12 · Nilo Pedrazzini

This paper explores the possibility of improving the performance of specialized parsers for pre-modern Slavic by training them on data from different related varieties. Because of their linguistic heterogeneity, pre-mode…

Dependency ParsingPart-Of-Speech TaggingPOSPOS Tagging

Unsupervised Bilingual Lexicon Induction Across Writing Systems

2020-01-31 · Parker Riley, Daniel Gildea

Recent embedding-based methods in unsupervised bilingual lexicon induction have shown good results, but generally have not leveraged orthographic (spelling) information, which can be helpful for pairs of related language…

Bilingual Lexicon Induction

Morphological Neural Pre- and Post-Processing for Slavic Languages

2019-08-01 · WS 2019 8 · Giorgio Bernardinello