paper-with-me

홈 › Papers

Towards a Better Understanding of Noise in Natural Language Processing

2021-09-01 · RANLP 2021 9 · Khetam Al Sharou, Zhenhao Li, Lucia Specia

In this paper, we propose a definition and taxonomy of various types of non-standard textual content – generally referred to as “noise” – in Natural Language Processing (NLP). While data pre-processing is undoubtedly important in NLP, especially when dealing with user-generated content, a broader understanding of different sources of noise and how to deal with them is an aspect that has been largely neglected. We provide a comprehensive list of potential sources of noise, categorise and describe them, and show the impact of a subset of standard pre-processing strategies on different tasks. Our main goal is to raise awareness of non-standard content – which should not always be considered as “noise” – and of the need for careful, task-dependent pre-processing. This is an alternative to blanket, all-encompassing solutions generally applied by researchers through “standard” pre-processing pipelines. The intention is for this categorisation to serve as a point of reference to support NLP researchers in devising strategies to clean, normalise or embrace non-standard content.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Joint Energy-based Model Training for Better Calibrated Natural Language Understanding Models

2021-01-18 · EACL 2021 2 · Tianxing He, Bryan McCann, Caiming Xiong, Ehsan Hosseini-Asl

In this work, we explore joint energy-based model (EBM) training during the finetuning of pretrained text encoders (e.g., Roberta) for natural language understanding (NLU) tasks. Our experiments show that EBM training ca…

Language ModelingLanguage ModellingNatural Language Understanding

An Approach to Improve Robustness of NLP Systems against ASR Errors

2021-03-25 · Tong Cui, Jinghui Xiao, Liangyou Li, Xin Jiang 외

Speech-enabled systems typically first convert audio to text through an automatic speech recognition (ASR) model and then feed the text to downstream natural language processing (NLP) modules. The errors of the ASR syste…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+5

Rethinking Denoised Auto-Encoding in Language Pre-Training

2021-11-01 · EMNLP 2021 11 · Fuli Luo, Pengcheng Yang, Shicheng Li, Xuancheng Ren 외

Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing. These models typically corrupt the given sequences with cer…

Natural Language UnderstandingSentence

CAPT: Contrastive Pre-Training for Learning Denoised Sequence Representations

2020-10-13 · Fuli Luo, Pengcheng Yang, Shicheng Li, Xuancheng Ren 외

Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing. These models typically corrupt the given sequences with cer…

Natural Language UnderstandingSentence

Towards Best Practices for Leveraging Human Language Processing Signals for Natural Language Processing

2020-05-01 · LREC 2020 5 · Nora Hollenstein, Maria Barrett, Lisa Beinborn

NLP models are imperfect and lack intricate capabilities that humans access automatically when processing speech or reading a text. Human language processing data can be leveraged to increase the performance of models an…

EEGElectroencephalogram (EEG)