paper-with-me

홈 › Papers

Revisiting NMT for Normalization of Early English Letters

2019-06-01 · WS 2019 6 · Mika H{\"a}m{\"a}l{\"a}inen, Tanja S{\"a}ily, Jack Rueter, J{\"o}rg Tiedemann, Eetu M{\"a}kel{\"a}

This paper studies the use of NMT (neural machine translation) as a normalization method for an early English letter corpus. The corpus has previously been normalized so that only less frequent deviant forms are left out without normalization. This paper discusses different methods for improving the normalization of these deviant forms by using different approaches. Adding features to the training data is found to be unhelpful, but using a lexicographical resource to filter the top candidates produced by the NMT model together with lemmatization improves results.

📄 PDF Abstract BibTeX

Code (1)

mikahama/natas

Tasks

LemmatizationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Normalizing Early English Letters to Present-day English Spelling

2018-08-01 · COLING 2018 8 · Mika H{\"a}m{\"a}l{\"a}inen, Tanja S{\"a}ily, Jack Rueter, J{\"o}rg Tiedemann 외

This paper presents multiple methods for normalizing the most deviant and infrequent historical spellings in a corpus consisting of personal correspondence from the 15th to the 19th century. The methods include machine t…

Machine TranslationTranslation

From Plenipotentiary to Puddingless: Users and Uses of New Words in Early English Letters

2021-03-17 · Tanja Säily, Eetu Mäkelä, Mika Hämäläinen

We study neologism use in two samples of early English correspondence, from 1640--1660 and 1760--1780. Of especial interest are the early adopters of new vocabulary, the social groups they represent, and the types and fu…

Graphemic Normalization of the Perso-Arabic Script

2022-10-21 · Raiomond Doctor, Alexander Gutkin, Cibu Johny, Brian Roark 외

Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various dia…

Language ModelingLanguage ModellingMachine Translation

The Dependence of Frequency Distributions on Multiple Meanings of Words, Codes and Signs

2017-09-28 · Xiao-Yong Yan, Petter Minnhagen

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This …

Part-of-Speech Tagging for Historical English

2016-03-10 · NAACL 2016 6 · Yi Yang, Jacob Eisenstein

As more historical texts are digitized, there is interest in applying natural language processing tools to these archives. However, the performance of these tools is often unsatisfactory, due to language change and genre…

Domain AdaptationPart-Of-Speech TaggingUnsupervised Domain AdaptationWord Embeddings