paper-with-me

홈 › Papers

Automatic Recognition of Linguistic Replacements in Text Series Generated from Keystroke Logs

2016-05-01 · LREC 2016 5 · Daniel Couto-Vale, Stella Neumann, Paula Niemietz

This paper introduces a toolkit used for the purpose of detecting replacements of different grammatical and semantic structures in ongoing text production logged as a chronological series of computer interaction events (so-called keystroke logs). The specific case we use involves human translations where replacements can be indicative of translator behaviour that leads to specific features of translations that distinguish them from non-translated texts. The toolkit uses a novel CCG chart parser customised so as to recognise grammatical words independently of space and punctuation boundaries. On the basis of the linguistic analysis, structures in different versions of the target text are compared and classified as potential equivalents of the same source text segment by {`}equivalence judges{'}. In that way, replacements of grammatical and semantic structures can be detected. Beyond the specific task at hand the approach will also be useful for the analysis of other types of spaceless text such as Twitter hashtags and texts in agglutinative or spaceless languages like Finnish or Chinese.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Investigating Lexical Replacements for Arabic-English Code-Switched Data Augmentation

2022-05-25 · Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu

Data sparsity is a main problem hindering the development of code-switching (CS) NLP systems. In this paper, we investigate data augmentation techniques for synthesizing dialectal Arabic-English CS text. We perform lexic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modelling+4

The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR

2025-03-30 · Injy Hamed, Ngoc Thang Vu, Nizar Habash

Code-switching, the act of alternating between languages, emerged as a prevalent global phenomenon that needs to be addressed for building user-friendly language technologies. A main bottleneck in this pursuit is data sc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translation+3

How Fragile is Relation Extraction under Entity Replacements?

2023-05-22 · Yiwei Wang, Bryan Hooi, Fei Wang, Yujun Cai 외

Relation extraction (RE) aims to extract the relations between entity names from the textual context. In principle, textual context determines the ground-truth relation and the RE models should be able to correctly ident…

BenchmarkingCausal InferenceRelationRelation Extraction

Context-sensitive evaluation of automatic speech recognition: considering user experience & language variation

2021-04-01 · EACL (HCINLP) 2021 4 · Nina Markl, Catherine Lai

Commercial Automatic Speech Recognition (ASR) systems tend to show systemic predictive bias for marginalised speaker/user groups. We highlight the need for an interdisciplinary and context-sensitive approach to documenti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Data Augmentation Techniques for Machine Translation of Code-Switched Texts: A Comparative Study

2023-10-23 · Injy Hamed, Nizar Habash, Ngoc Thang Vu

Code-switching (CSW) text generation has been receiving increasing attention as a solution to address data scarcity. In light of this growing interest, we need more comprehensive studies comparing different augmentation …

Data AugmentationMachine TranslationText GenerationTranslation