Improving historical spelling normalization with bi-directional LSTMs and multi-task learning
Natural-language processing of historical documents is complicated by the abundance of variant spellings and lack of annotated data. A common approach is to normalize the spelling of historical words to modern forms. We explore the suitability of a deep neural network architecture for this task, particularly a deep bi-LSTM network applied on a character level. Our model compares well to previously established normalization algorithms when evaluated on a diverse set of texts from Early New High German. We show that multi-task learning with additional normalization data can improve our model's performance further.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Task LearningSimilar Papers 제목 키워드 기반
An Evaluation of Neural Machine Translation Models on Historical Spelling Normalization
In this paper, we apply different NMT models to the problem of historical spelling normalization for five languages: English, German, Hungarian, Icelandic, and Swedish. The NMT models are at different levels, have differ…
Machine TranslationNMTTranslationEvaluating Inter-Annotator Agreement on Historical Spelling Normalization
Historical German Text Normalization Using Type- and Token-Based Language Modeling
Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually…
DecoderLanguage ModelingLanguage ModellingLarge Language Model+2Using Comparable Collections of Historical Texts for Building a Diachronic Dictionary for Spelling Normalization
Normalizing Early English Letters to Present-day English Spelling
This paper presents multiple methods for normalizing the most deviant and infrequent historical spellings in a corpus consisting of personal correspondence from the 15th to the 19th century. The methods include machine t…
Machine TranslationTranslation