Identifying Cognate Sets Across Dictionaries of Related Languages
We present a system for identifying cognate sets across dictionaries of related languages. The likelihood of a cognate relationship is calculated on the basis of a rich set of features that capture both phonetic and semantic similarity, as well as the presence of regular sound correspondences. The similarity scores are used to cluster words from different languages that may originate from a common proto-word. When tested on the Algonquian language family, our system detects 63{\%} of cognate sets while maintaining cluster purity of 70{\%}.
Code (1)
Tasks
Semantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Building a Dataset of Multilingual Cognates for the Romanian Lexicon
Identifying cognates is an interesting task with applications in numerous research areas, such as historical and comparative linguistics, language acquisition, cross-lingual information retrieval, readability and machine…
Cross-Lingual Information RetrievalInformation RetrievalLanguage AcquisitionMachine Translation+3Automatic Identification and Production of Related Words for Historical Linguistics
Language change across space and time is one of the main concerns in historical linguistics. In this article, we develop tools to assist researchers and domain experts in the study of language evolution.First, we introdu…
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer
Identification of cognates across related languages is one of the primary problems in historical linguistics. Automated cognate identification is helpful for several downstream tasks including identifying sound correspon…
Link PredictionBilingual Lexicon Induction across Orthographically-distinct Under-Resourced Dravidian Languages
Bilingual lexicons are a vital tool for under-resourced languages and recent state-of-the-art approaches to this leverage pretrained monolingual word embeddings using supervised or semi-supervised approaches. However, th…
Bilingual Lexicon InductionWord EmbeddingsA Generalized Constraint Approach to Bilingual Dictionary Induction for Low-Resource Language Families
The lack or absence of parallel and comparable corpora makes bilingual lexicon extraction a difficult task for low-resource languages. The pivot language and cognate recognition approaches have been proven useful for ind…
Bilingual Lexicon InductionTranslationWord Alignment