paper-with-me

홈 › Papers

Comparing Fifty Natural Languages and Twelve Genetic Languages Using Word Embedding Language Divergence (WELD) as a Quantitative Measure of Language Distance

2016-04-28 · WS 2016 6 · Ehsaneddin Asgari, Mohammad R. K. Mofrad

We introduce a new measure of distance between languages based on word embedding, called word embedding language divergence (WELD). WELD is defined as divergence between unified similarity distribution of words between languages. Using such a measure, we perform language comparison for fifty natural languages and twelve genetic languages. Our natural language dataset is a collection of sentence-aligned parallel corpora from bible translations for fifty languages spanning a variety of language families. Although we use parallel corpora, which guarantees having the same content in all languages, interestingly in many cases languages within the same family cluster together. In addition to natural languages, we perform language comparison for the coding regions in the genomes of 12 different organisms (4 plants, 6 animals, and two human subjects). Our result confirms a significant high-level difference in the genetic language model of humans/animals versus plants. The proposed method is a step toward defining a quantitative measure of similarity between languages, with applications in languages classification, genre identification, dialect identification, and evaluation of translations.

📄 PDF Abstract BibTeX arXiv:1604.08561

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect IdentificationLanguage ModelingLanguage ModellingSentence

Similar Papers 제목 키워드 기반

Challenge Dataset of Cognates and False Friend Pairs from Indian Languages

2021-12-17 · LREC 2020 5 · Diptesh Kanojia, Pushpak Bhattacharyya, Malhar Kulkarni, Gholamreza Haffari

Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP…

Information RetrievalMachine TranslationRetrievalTranslation

Human genetic admixture through the lens of population genomics

2021-09-24 · Shyamalika Gopalan, Samuel Patillo Smith, Katharine Korunes, Iman Hamid 외

Over the last fifty years, geneticists have made great strides in understanding how our species' evolutionary history gave rise to current patterns of human genetic diversity classically summarized by Lewontin in his 197…

Diversity

Measuring the Similarity of Grammatical Gender Systems by Comparing Partitions

2020-11-01 · EMNLP 2020 11 · Arya D. McCarthy, Adina Williams, Shijia Liu, David Yarowsky 외

A grammatical gender system divides a lexicon into a small number of relatively fixed grammatical categories. How similar are these gender systems across languages? To quantify the similarity, we define gender systems ex…

Community Detection

What do character-level models learn about morphology? The case of dependency parsing

2018-08-28 · EMNLP 2018 10 · Clara Vania, Andreas Grivas, Adam Lopez

When parsing morphologically-rich languages with neural models, it is beneficial to model input at the character level, and it has been claimed that this is because character-level models learn morphology. We test these …

Dependency ParsingMorphological Analysis

Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages

2021-12-16 · COLING 2020 8 · Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya 외

Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of c…

Cross-Lingual Information RetrievalCross-Lingual Word EmbeddingsInformation RetrievalMachine Translation+4