Punctuation as Native Language Interference
In this paper, we describe experiments designed to explore and evaluate the impact of punctuation marks on the task of native language identification. Punctuation is specific to each language, and is part of the indicators that overtly represent the manner in which each language organizes and conveys information. Our experiments are organized in various set-ups: the usual multi-class classification for individual languages, also considering classification by language groups, across different proficiency levels, topics and even cross-corpus. The results support our hypothesis that punctuation marks are persistent and robust indicators of the native language of the author, which do not diminish in influence even when a high proficiency level in a non-native language is achieved.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationCross-corpusGeneral ClassificationLanguage IdentificationMulti-class ClassificationNative Language IdentificationSimilar Papers 제목 키워드 기반
Nightmare at test time: How punctuation prevents parsers from generalizing
Punctuation is a strong indicator of syntactic structure, and parsers trained on text with punctuation often rely heavily on this signal. Punctuation is a diversion, however, since human language processing does not rely…
Discriminative Self-training for Punctuation Prediction
Punctuation prediction for automatic speech recognition (ASR) output transcripts plays a crucial role for improving the readability of the ASR transcripts and for improving the performance of downstream natural language …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Predictionspeech-recognition+1PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text
Text transcripts without punctuation or sentence boundaries are hard to comprehend for both humans and machines. Punctuation marks play a vital role by providing meaning to the sentence and incorrect use or placement of …
Language ModellingPunctuation RestorationSentenceTranslationToken-Level Supervised Contrastive Learning for Punctuation Restoration
Punctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent de…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningIntent Detection+5Comprehensive Punctuation Restoration for English and Polish
Punctuation restoration is a fundamental requirement for the readability of text derived from Automatic Speech Recognition (ASR) systems. Most contemporary solutions are limited to predicting only a few of the most frequ…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Punctuation Restorationspeech-recognition+1