Revisiting NMT for Normalization of Early English Letters
This paper studies the use of NMT (neural machine translation) as a normalization method for an early English letter corpus. The corpus has previously been normalized so that only less frequent deviant forms are left out without normalization. This paper discusses different methods for improving the normalization of these deviant forms by using different approaches. Adding features to the training data is found to be unhelpful, but using a lexicographical resource to filter the top candidates produced by the NMT model together with lemmatization improves results.
Code (1)
Tasks
LemmatizationMachine TranslationNMTTranslationSimilar Papers 제목 키워드 기반
Normalizing Early English Letters to Present-day English Spelling
This paper presents multiple methods for normalizing the most deviant and infrequent historical spellings in a corpus consisting of personal correspondence from the 15th to the 19th century. The methods include machine t…
Machine TranslationTranslationFrom Plenipotentiary to Puddingless: Users and Uses of New Words in Early English Letters
We study neologism use in two samples of early English correspondence, from 1640--1660 and 1760--1780. Of especial interest are the early adopters of new vocabulary, the social groups they represent, and the types and fu…
Graphemic Normalization of the Perso-Arabic Script
Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various dia…
Language ModelingLanguage ModellingMachine TranslationThe Dependence of Frequency Distributions on Multiple Meanings of Words, Codes and Signs
The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This …
Part-of-Speech Tagging for Historical English
As more historical texts are digitized, there is interest in applying natural language processing tools to these archives. However, the performance of these tools is often unsatisfactory, due to language change and genre…
Domain AdaptationPart-Of-Speech TaggingUnsupervised Domain AdaptationWord Embeddings