Probing Multilingual Cognate Prediction Models
Character-based neural machine translation models have become the reference models for cognate prediction, a historical linguistics task. So far, all linguistic interpretations about latent information captured by such models have been based on external analysis (accuracy, raw results, errors). In this paper, we investigate what probing can tell us about both models and previous interpretations, and learn that though our models store linguistic and diachronic information, they do not achieve it in previously assumed ways.
Code (0)
등록된 구현이 없습니다.
Tasks
Cognate PredictionMachine TranslationPredictionTranslationSimilar Papers 제목 키워드 기반
The SIGTYP 2022 Shared Task on the Prediction of Cognate Reflexes
This study describes the structure and the results of the SIGTYP 2022 shared task on the prediction of cognate reflexes from multilingual wordlists. We asked participants to submit systems that would predict words in ind…
Image RestorationAutomated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer
Identification of cognates across related languages is one of the primary problems in historical linguistics. Automated cognate identification is helpful for several downstream tasks including identifying sound correspon…
Link PredictionSimilarity Dependent Chinese Restaurant Process for Cognate Identification in Multilingual Wordlists
We present and evaluate two similarity dependent Chinese Restaurant Process (sd-CRP) algorithms at the task of automated cognate detection. The sd-CRP clustering algorithms do not require any predefined threshold for det…
ClusteringLanguage IdentificationCognate-aware morphological segmentation for multilingual neural translation
This article describes the Aalto University entry to the WMT18 News Translation Shared Task. We participate in the multilingual subtrack with a system trained under the constrained condition to translate from English to …
TranslationPhonetic, Semantic, and Articulatory Features in Assamese-Bengali Cognate Detection
In this paper, we propose a method to detect if words in two similar languages, Assamese and Bengali, are cognates. We mix phonetic, semantic, and articulatory features and use the cognate detection task to analyze the r…