paper-with-me

홈 › Papers

Identifying Cognate Sets Across Dictionaries of Related Languages

2017-09-01 · EMNLP 2017 9 · Adam St Arnaud, David Beck, Grzegorz Kondrak

We present a system for identifying cognate sets across dictionaries of related languages. The likelihood of a cognate relationship is calculated on the basis of a rich set of features that capture both phonetic and semantic similarity, as well as the presence of regular sound correspondences. The similarity scores are used to cluster words from different languages that may originate from a common proto-word. When tested on the Algonquian language family, our system detects 63{\%} of cognate sets while maintaining cluster purity of 70{\%}.

📄 PDF Abstract BibTeX

Code (1)

ajstarna/SemaPhoR 공식 구현

Tasks

Semantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Building a Dataset of Multilingual Cognates for the Romanian Lexicon

2014-05-01 · LREC 2014 5 · Liviu Dinu, Alina Maria Ciobanu

Identifying cognates is an interesting task with applications in numerous research areas, such as historical and comparative linguistics, language acquisition, cross-lingual information retrieval, readability and machine…

Cross-Lingual Information RetrievalInformation RetrievalLanguage AcquisitionMachine Translation+3

Automatic Identification and Production of Related Words for Historical Linguistics

2019-12-01 · CL 2019 12 · Alina Maria Ciobanu, Liviu P. Dinu

Language change across space and time is one of the main concerns in historical linguistics. In this article, we develop tools to assist researchers and domain experts in the study of language evolution.First, we introdu…

Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer

2024-02-05 · V. S. D. S. Mahesh Akavarapu, Arnab Bhattacharya

Identification of cognates across related languages is one of the primary problems in historical linguistics. Automated cognate identification is helpful for several downstream tasks including identifying sound correspon…

Link Prediction

Bilingual Lexicon Induction across Orthographically-distinct Under-Resourced Dravidian Languages

2020-12-01 · VarDial (COLING) 2020 12 · Bharathi Raja Chakravarthi, Navaneethan Rajasekaran, Mihael Arcan, Kevin McGuinness 외

Bilingual lexicons are a vital tool for under-resourced languages and recent state-of-the-art approaches to this leverage pretrained monolingual word embeddings using supervised or semi-supervised approaches. However, th…

Bilingual Lexicon InductionWord Embeddings

A Generalized Constraint Approach to Bilingual Dictionary Induction for Low-Resource Language Families

2020-10-05 · Arbi Haza Nasution, Yohei Murakami, Toru Ishida

The lack or absence of parallel and comparable corpora makes bilingual lexicon extraction a difficult task for low-resource languages. The pivot language and cognate recognition approaches have been proven useful for ind…

Bilingual Lexicon InductionTranslationWord Alignment