paper-with-me

Papers

Challenge Dataset of Cognates and False Friend Pairs from Indian Languages

2021-12-17 · LREC 2020 5 · Diptesh Kanojia, Pushpak Bhattacharyya, Malhar Kulkarni, Gholamreza Haffari

Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP) applications such as Machine Translation, Cross-lingual Sense Disambiguation, Computational Phylogenetics, and Information Retrieval. A possible solution to address this challenge is to identify cognates across language pairs. In this paper, we describe the creation of two cognate datasets for twelve Indian languages, namely Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam. We digitize the cognate data from an Indian language cognate dictionary and utilize linked Indian language Wordnets to generate cognate sets. Additionally, we use the Wordnet data to create a False Friends' dataset for eleven language pairs. We also evaluate the efficacy of our dataset using previously available baseline cognate detection approaches. We also perform a manual evaluation with the help of lexicographers and release the curated gold-standard dataset with this paper.

📄 PDF Abstract BibTeX arXiv:2112.09526

Code (1)

dipteshkanojia/challengeCognateFF 공식 구현

Tasks

Information RetrievalMachine TranslationRetrievalTranslation

Similar Papers 제목 키워드 기반

Automatically Building a Multilingual Lexicon of False Friends With No Supervision

2020-05-01 · LREC 2020 5 · Ana Sabina Uban, Liviu P. Dinu

Cognate words, defined as words in different languages which derive from a common etymon, can be useful for language learners, who can leverage the orthographical similarity of cognates to more easily understand a text i…

Cross-Lingual Word EmbeddingsLanguage AcquisitionWord Embeddings

A Computational Approach to Measuring the Semantic Divergence of Cognates

2020-12-02 · Ana-Sabina Uban, Alina-Maria Ciobanu, Liviu P. Dinu

Meaning is the foundation stone of intercultural communication. Languages are continuously changing, and words shift their meanings for various reasons. Semantic divergence in related languages is a key concern of histor…

Cross-Lingual Word EmbeddingsSemantic SimilaritySemantic Textual SimilarityTranslation+1

Tracking Semantic Change in Cognate Sets for English and Romance Languages

2021-08-01 · ACL (LChange) 2021 8 · Ana Sabina Uban, Alina Maria Cristea, Anca Dinu, Liviu P. Dinu 외

Semantic divergence in related languages is a key concern of historical linguistics. We cross-linguistically investigate the semantic divergence of cognate pairs in English and Romance languages, by means of word embeddi…

Word Embeddings

When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates

2026-06-11 · Junhong Liang, Noor Abo Mokh, Bashar Alhafni arxiv

Arabic and Hebrew, as closely related Semitic languages, share a substantial lexicon of true cognates, misleading false friends, and modern loanwords. This overlap poses a challenge for cross-lingual semantic understandi…

A Classification-Based Approach to Cognate Detection Combining Orthographic and Semantic Similarity Information

2019-09-01 · RANLP 2019 9 · Sofie Labat, Els Lefever

This paper presents proof-of-concept experiments for combining orthographic and semantic information to distinguish cognates from non-cognates. To this end, a context-independent gold standard is developed by manually la…

Binary ClassificationFormGeneral ClassificationSemantic Similarity+2