paper-with-me

Papers

From Isolates to Families: Using Neural Networks for Automated Language Affiliation

2025-02-17 · Frederic Blum, Steffen Herbold, Johann-Mattis List

In historical linguistics, the affiliation of languages to a common language family is traditionally carried out using a complex workflow that relies on manually comparing individual languages. Large-scale standardized collections of multilingual wordlists and grammatical language structures might help to improve this and open new avenues for developing automated language affiliation workflows. Here, we present neural network models that use lexical and grammatical data from a worldwide sample of more than 1,000 languages with known affiliations to classify individual languages into families. In line with the traditional assumption of most linguists, our results show that models trained on lexical data alone outperform models solely based on grammatical data, whereas combining both types of data yields even better performance. In additional experiments, we show how our models can identify long-ranging relations between entire subgroups, how they can be employed to investigate potential relatives of linguistic isolates, and how they can help us to obtain first hints on the affiliation of so far unaffiliated languages. We conclude that models for automated language affiliation trained on lexical and grammatical data provide comparative linguists with a valuable tool for evaluating hypotheses about deep and unknown language relations.

📄 PDF Abstract BibTeX arXiv:2502.11688

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-Linguistic Examination of Machine Translation Transfer Learning

2024-12-27 · Saughmon Boujkian

This study investigates the effectiveness of transfer learning in machine translation across diverse linguistic families by evaluating five distinct language pairs. Leveraging pre-trained models on high-resource language…

Machine TranslationTransfer LearningTranslation

Phylogenetic typology

2021-03-18 · Gerhard Jäger, Johannes Wahle

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our techniqu…

Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers

2020-08-01 · WOSP 2020 8 · Yoshitomo Matsubara, Sameer Singh

The question of the utility of the blind peer-review system is fundamental to scientific research. Some studies investigate exactly how “blind” the papers are in the double-blind review system by manually or automaticall…

Word length predicts word order: "Min-max"-ing drives language evolution

2025-05-20 · Hiram Ring

Current theories of language propose an innate (Baker 2001; Chomsky 1981) or a functional (Greenberg 1963; Dryer 2007; Hawkins 2014) origin for the surface structures (i.e. word order) that we observe in languages of the…

From raw affiliations to organization identifiers

2025-05-12 · Myrto Kallipoliti, Serafeim Chatzopoulos, Miriam Baglioni, Eleni Adamidi 외

Accurate affiliation matching, which links affiliation strings to standardized organization identifiers, is critical for improving research metadata quality, facilitating comprehensive bibliometric analyses, and supporti…

BenchmarkingMetadata quality