The Cognate Data Bottleneck in Language Phylogenetics
To fully exploit the potential of computational phylogenetic methods for cognate data one needs to leverage specific (complex) models an machine learning-based techniques. However, both approaches require datasets that are substantially larger than the manually collected cognate data currently available. To the best of our knowledge, there exists no feasible approach to automatically generate larger cognate datasets. We substantiate this claim by automatically extracting datasets from BabelNet, a large multilingual encyclopedic dictionary. We demonstrate that phylogenetic inferences on the respective character matrices yield trees that are largely inconsistent with the established gold standard ground truth trees. We also discuss why we consider it as being unlikely to be able to extract more suitable character matrices from other multilingual resources. Phylogenetic data analysis approaches that require larger datasets can therefore not be applied to cognate data. Thus, it remains an open question how, and if these computational approaches can be applied in historical linguistics.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Beyond cognacy
Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-based inference. However, standard approaches rely on expert-annotated cogna…
Multiple Sequence AlignmentChallenge Dataset of Cognates and False Friend Pairs from Indian Languages
Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP…
Information RetrievalMachine TranslationRetrievalTranslationUtilizing Wordnets for Cognate Detection among Indian Languages
Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics. Unidentified cognate pairs can pos…
Information RetrievalMachine TranslationRetrievalHarnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages
Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of c…
Cross-Lingual Information RetrievalCross-Lingual Word EmbeddingsInformation RetrievalMachine Translation+4Cognition-aware Cognate Detection
Automatic detection of cognates helps downstream NLP tasks of Machine Translation, Cross-lingual Information Retrieval, Computational Phylogenetics and Cross-lingual Named Entity Recognition. Previous approaches for the …
Cross-Lingual Information RetrievalInformation RetrievalMachine Translationnamed-entity-recognition+6