paper-with-me

홈 › Papers

The Cognate Data Bottleneck in Language Phylogenetics

2025-07-01 · Luise Häuser, Alexandros Stamatakis arxiv

To fully exploit the potential of computational phylogenetic methods for cognate data one needs to leverage specific (complex) models an machine learning-based techniques. However, both approaches require datasets that are substantially larger than the manually collected cognate data currently available. To the best of our knowledge, there exists no feasible approach to automatically generate larger cognate datasets. We substantiate this claim by automatically extracting datasets from BabelNet, a large multilingual encyclopedic dictionary. We demonstrate that phylogenetic inferences on the respective character matrices yield trees that are largely inconsistent with the established gold standard ground truth trees. We also discuss why we consider it as being unlikely to be able to extract more suitable character matrices from other multilingual resources. Phylogenetic data analysis approaches that require larger datasets can therefore not be applied to cognate data. Thus, it remains an open question how, and if these computational approaches can be applied in historical linguistics.

📄 PDF Abstract BibTeX arXiv:2507.00911

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond cognacy

2025-07-02 · Gerhard Jäger arxiv

Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-based inference. However, standard approaches rely on expert-annotated cogna…

Multiple Sequence Alignment

Challenge Dataset of Cognates and False Friend Pairs from Indian Languages

2021-12-17 · LREC 2020 5 · Diptesh Kanojia, Pushpak Bhattacharyya, Malhar Kulkarni, Gholamreza Haffari

Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP…

Information RetrievalMachine TranslationRetrievalTranslation

Utilizing Wordnets for Cognate Detection among Indian Languages

2021-12-30 · GWC 2019 7 · Diptesh Kanojia, Kevin Patel, Pushpak Bhattacharyya, Malhar Kulkarni 외

Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics. Unidentified cognate pairs can pos…

Information RetrievalMachine TranslationRetrieval

Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

2021-12-16 · COLING 2020 8 · Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya 외

Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of c…

Cross-Lingual Information RetrievalCross-Lingual Word EmbeddingsInformation RetrievalMachine Translation+4

Cognition-aware Cognate Detection

2021-12-15 · EACL 2021 2 · Diptesh Kanojia, Prashant Sharma, Sayali Ghodekar, Pushpak Bhattacharyya 외

Automatic detection of cognates helps downstream NLP tasks of Machine Translation, Cross-lingual Information Retrieval, Computational Phylogenetics and Cross-lingual Named Entity Recognition. Previous approaches for the …

Cross-Lingual Information RetrievalInformation RetrievalMachine Translationnamed-entity-recognition+6