paper-with-me

Papers

A Classification-Based Approach to Cognate Detection Combining Orthographic and Semantic Similarity Information

2019-09-01 · RANLP 2019 9 · Sofie Labat, Els Lefever

This paper presents proof-of-concept experiments for combining orthographic and semantic information to distinguish cognates from non-cognates. To this end, a context-independent gold standard is developed by manually labelling English-Dutch pairs of cognates and false friends in bilingual term lists. These annotated cognate pairs are then used to train and evaluate a supervised binary classification system for the automatic detection of cognates. Two types of information sources are incorporated in the classifier: fifteen string similarity metrics capture form similarity between source and target words, while word embeddings model semantic similarity between the words. The experimental results show that even though the system already achieves good results by only incorporating orthographic information, the performance further improves by including semantic information in the form of embeddings.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationFormGeneral ClassificationSemantic SimilaritySemantic Textual SimilarityWord Embeddings

Similar Papers 제목 키워드 기반

Identifying Cognates in English-Dutch and French-Dutch by means of Orthographic Information and Cross-lingual Word Embeddings

2020-05-01 · LREC 2020 5 · Els Lefever, Sofie Labat, Pranaydeep Singh

This paper investigates the validity of combining more traditional orthographic information with cross-lingual word embeddings to identify cognate pairs in English-Dutch and French-Dutch. In a first step, lists of potent…

Cross-Lingual Word EmbeddingsWord Embeddings

Automatic Detection of Cognates Using Orthographic Alignment

2014-06-01 · ACL 2014 6 · Alina Maria Ciobanu, Liviu P. Dinu
Information RetrievalLanguage AcquisitionMachine TranslationSemantic Textual Similarity

Multilingual LLMs Struggle to Link Orthography and Semantics in Bilingual Word Processing

2025-01-15 · Eshaan Tanwar, Gayatri Oke, Tanmoy Chakraborty

Bilingual lexical processing is shaped by the complex interplay of phonological, orthographic, and semantic features of two languages within an integrated mental lexicon. In humans, this is evident in the ease with which…

Cognition-aware Cognate Detection

2021-12-15 · EACL 2021 2 · Diptesh Kanojia, Prashant Sharma, Sayali Ghodekar, Pushpak Bhattacharyya 외

Automatic detection of cognates helps downstream NLP tasks of Machine Translation, Cross-lingual Information Retrieval, Computational Phylogenetics and Cross-lingual Named Entity Recognition. Previous approaches for the …

Cross-Lingual Information RetrievalInformation RetrievalMachine Translationnamed-entity-recognition+6

Building a Dataset of Multilingual Cognates for the Romanian Lexicon

2014-05-01 · LREC 2014 5 · Liviu Dinu, Alina Maria Ciobanu

Identifying cognates is an interesting task with applications in numerous research areas, such as historical and comparative linguistics, language acquisition, cross-lingual information retrieval, readability and machine…

Cross-Lingual Information RetrievalInformation RetrievalLanguage AcquisitionMachine Translation+3