paper-with-me

Papers

USZEGED: Correction Type-sensitive Normalization of English Tweets Using Efficiently Indexed n-gram Statistics

2015-07-01 · WS 2015 7 · G{\'a}bor Berend, Ervin Tasn{\'a}di
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Lexical Normalization

Similar Papers 제목 키워드 기반

Visual Cues and Error Correction for Translation Robustness

2021-03-12 · Findings (EMNLP) 2021 11 · Zhenhao Li, Marek Rei, Lucia Specia

Neural Machine Translation models are sensitive to noise in the input texts, such as misspelled words and ungrammatical constructions. Existing robustness techniques generally fail when faced with unseen types of noise a…

Machine TranslationTranslation

MoNoise: Modeling Noise Using a Modular Normalization System

2017-10-10 · Rob van der Goot, Gertjan van Noord

We propose MoNoise: a normalization model focused on generalizability and efficiency, it aims at being easily reusable and adaptable. Normalization is the task of translating texts from a non- canonical domain to a more …

Lexical NormalizationSpelling CorrectionWord Embeddings

Unsupervised Context-Sensitive Spelling Correction of English and Dutch Clinical Free-Text with Word and Character N-Gram Embeddings

2017-10-19 · Pieter Fivez, Simon Šuster, Walter Daelemans

We present an unsupervised context-sensitive spelling correction method for clinical free-text that uses word and character n-gram embeddings. Our method generates misspelling replacement candidates and ranks them accord…

Spelling Correction

TNT: Text Normalization based Pre-training of Transformers for Content Moderation

2020-11-01 · EMNLP 2020 11 · Fei Tan, Yifan Hu, Changwei Hu, Keqian Li 외

In this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation. Inspired by the masking strategy and text normalization, TNT is developed to …

Text Normalization

IDENTIC Corpus: Morphologically Enriched Indonesian-English Parallel Corpus

2012-05-01 · LREC 2012 5 · Septina Dian Larasati

This paper describes the creation process of an Indonesian-English parallel corpus (IDENTIC). The corpus contains 45,000 sentences collected from different sources in different genres. Several manual text preprocessing t…

Spelling Correction