Towards a Better Understanding of Noise in Natural Language Processing
In this paper, we propose a definition and taxonomy of various types of non-standard textual content – generally referred to as “noise” – in Natural Language Processing (NLP). While data pre-processing is undoubtedly important in NLP, especially when dealing with user-generated content, a broader understanding of different sources of noise and how to deal with them is an aspect that has been largely neglected. We provide a comprehensive list of potential sources of noise, categorise and describe them, and show the impact of a subset of standard pre-processing strategies on different tasks. Our main goal is to raise awareness of non-standard content – which should not always be considered as “noise” – and of the need for careful, task-dependent pre-processing. This is an alternative to blanket, all-encompassing solutions generally applied by researchers through “standard” pre-processing pipelines. The intention is for this categorisation to serve as a point of reference to support NLP researchers in devising strategies to clean, normalise or embrace non-standard content.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Joint Energy-based Model Training for Better Calibrated Natural Language Understanding Models
In this work, we explore joint energy-based model (EBM) training during the finetuning of pretrained text encoders (e.g., Roberta) for natural language understanding (NLU) tasks. Our experiments show that EBM training ca…
Language ModelingLanguage ModellingNatural Language UnderstandingAn Approach to Improve Robustness of NLP Systems against ASR Errors
Speech-enabled systems typically first convert audio to text through an automatic speech recognition (ASR) model and then feed the text to downstream natural language processing (NLP) modules. The errors of the ASR syste…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+5Rethinking Denoised Auto-Encoding in Language Pre-Training
Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing. These models typically corrupt the given sequences with cer…
Natural Language UnderstandingSentenceCAPT: Contrastive Pre-Training for Learning Denoised Sequence Representations
Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing. These models typically corrupt the given sequences with cer…
Natural Language UnderstandingSentenceTowards Best Practices for Leveraging Human Language Processing Signals for Natural Language Processing
NLP models are imperfect and lack intricate capabilities that humans access automatically when processing speech or reading a text. Human language processing data can be leveraged to increase the performance of models an…
EEGElectroencephalogram (EEG)