Automatic Recognition of Linguistic Replacements in Text Series Generated from Keystroke Logs
This paper introduces a toolkit used for the purpose of detecting replacements of different grammatical and semantic structures in ongoing text production logged as a chronological series of computer interaction events (so-called keystroke logs). The specific case we use involves human translations where replacements can be indicative of translator behaviour that leads to specific features of translations that distinguish them from non-translated texts. The toolkit uses a novel CCG chart parser customised so as to recognise grammatical words independently of space and punctuation boundaries. On the basis of the linguistic analysis, structures in different versions of the target text are compared and classified as potential equivalents of the same source text segment by {`}equivalence judges{'}. In that way, replacements of grammatical and semantic structures can be detected. Beyond the specific task at hand the approach will also be useful for the analysis of other types of spaceless text such as Twitter hashtags and texts in agglutinative or spaceless languages like Finnish or Chinese.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Investigating Lexical Replacements for Arabic-English Code-Switched Data Augmentation
Data sparsity is a main problem hindering the development of code-switching (CS) NLP systems. In this paper, we investigate data augmentation techniques for synthesizing dialectal Arabic-English CS text. We perform lexic…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modelling+4The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR
Code-switching, the act of alternating between languages, emerged as a prevalent global phenomenon that needs to be addressed for building user-friendly language technologies. A main bottleneck in this pursuit is data sc…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translation+3How Fragile is Relation Extraction under Entity Replacements?
Relation extraction (RE) aims to extract the relations between entity names from the textual context. In principle, textual context determines the ground-truth relation and the RE models should be able to correctly ident…
BenchmarkingCausal InferenceRelationRelation ExtractionContext-sensitive evaluation of automatic speech recognition: considering user experience & language variation
Commercial Automatic Speech Recognition (ASR) systems tend to show systemic predictive bias for marginalised speaker/user groups. We highlight the need for an interdisciplinary and context-sensitive approach to documenti…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionData Augmentation Techniques for Machine Translation of Code-Switched Texts: A Comparative Study
Code-switching (CSW) text generation has been receiving increasing attention as a solution to address data scarcity. In light of this growing interest, we need more comprehensive studies comparing different augmentation …
Data AugmentationMachine TranslationText GenerationTranslation