A High Coverage Method for Automatic False Friends Detection for Spanish and Portuguese
False friends are words in two languages that look or sound similar, but have different meanings. They are a common source of confusion among language learners. Methods to detect them automatically do exist, however they make use of large aligned bilingual corpora, which are hard to find and expensive to build, or encounter problems dealing with infrequent words. In this work we propose a high coverage method that uses word vector representations to build a false friends classifier for any pair of languages, which we apply to the particular case of Spanish and Portuguese. The required resources are a large corpus for each language and a small bilingual lexicon for the pair.
Code (1)
Similar Papers 제목 키워드 기반
Automatically Building a Multilingual Lexicon of False Friends With No Supervision
Cognate words, defined as words in different languages which derive from a common etymon, can be useful for language learners, who can leverage the orthographical similarity of cognates to more easily understand a text i…
Cross-Lingual Word EmbeddingsLanguage AcquisitionWord EmbeddingsWidening the Discussion on ``False Friends'' in Multilingual Wordnets
There are wordnets in many languages, many aligned with Princeton WordNet, some of which in a (semi-)automatic process, but we rarely see actual discussions on the role of false friends in this process. Having in mind kn…
TranslationA Computational Approach to Measuring the Semantic Divergence of Cognates
Meaning is the foundation stone of intercultural communication. Languages are continuously changing, and words shift their meanings for various reasons. Semantic divergence in related languages is a key concern of histor…
Cross-Lingual Word EmbeddingsSemantic SimilaritySemantic Textual SimilarityTranslation+1Is Cross-Lingual Transfer in Bilingual Models Human-Like? A Study with Overlapping Word Forms in Dutch and English
Bilingual speakers show cross-lingual activation during reading, especially for words with shared surface form. Cognates (friends) typically lead to facilitation, whereas interlingual homographs (false friends) cause int…
Cross-Lingual TransferFalse-Friend Detection and Entity Matching via Unsupervised Transliteration
Transliterations play an important role in multilingual entity reference resolution, because proper names increasingly travel between languages in news and social media. Previous work associated with machine translation …
Machine TranslationTranslationTransliteration