Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics?
We evaluate the performance of state-of-the-art algorithms for automatic cognate detection by comparing how useful automatically inferred cognates are for the task of phylogenetic inference compared to classical manually annotated cognate sets. Our findings suggest that phylogenies inferred from automated cognate sets come close to phylogenies inferred from expert-annotated ones, although on average, the latter are still superior. We conclude that future work on phylogenetic reconstruction can profit much from automatic cognate detection. Especially where scholars are merely interested in exploring the bigger picture of a language family's phylogeny, algorithms for automatic cognate detection are a useful complement for current research on language phylogenies.
Code (1)
Similar Papers 제목 키워드 기반
A Classification-Based Approach to Cognate Detection Combining Orthographic and Semantic Similarity Information
This paper presents proof-of-concept experiments for combining orthographic and semantic information to distinguish cognates from non-cognates. To this end, a context-independent gold standard is developed by manually la…
Binary ClassificationFormGeneral ClassificationSemantic Similarity+2Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages
Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of c…
Cross-Lingual Information RetrievalCross-Lingual Word EmbeddingsInformation RetrievalMachine Translation+4Utilizing Wordnets for Cognate Detection among Indian Languages
Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics. Unidentified cognate pairs can pos…
Information RetrievalMachine TranslationRetrievalAlignment Analysis of Sequential Segmentation of Lexicons to Improve Automatic Cognate Detection
Ranking functions in information retrieval are often used in search engines to recommend the relevant answers to the query. This paper makes use of this notion of information retrieval and applies onto the problem domain…
General ClassificationInformation RetrievalLanguage ModellingRetrieval+1