paper-with-me

Papers

Pro-TEXT: an Annotated Corpus of Keystroke Logs

2022-06-01 · LREC 2022 6 · Aleksandra Miletic, Christophe Benzitoun, Georgeta Cislaru, Santiago Herrera-Yanez

Pro-TEXT is a corpus of keystroke logs written in French. Keystroke logs are recordings of the writing process executed through a keyboard, which keep track of all actions taken by the writer (character additions, deletions, substitutions). As such, the Pro-TEXT corpus offers new insights into text genesis and underlying cognitive processes from the production perspective. A subset of the corpus is linguistically annotated with parts of speech, lemmas and syntactic dependencies, making it suitable for the study of interactions between linguistic and behavioural aspects of the writing process. The full corpus contains 202K tokens, while the annotated portion is currently 30K tokens large. The annotated content is progressively being made available in a database-like CSV format and in CoNLL format, and the work on an HTML-based visualisation tool is currently under way. To the best of our knowledge, Pro-TEXT is the first corpus of its kind in French.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Keystroke dynamics as signal for shallow syntactic parsing

2016-10-11 · COLING 2016 12 · Barbara Plank

Keystroke dynamics have been extensively used in psycholinguistic and writing research to gain insights into cognitive processing. But do keystroke logs contain actual signal that can be used to learn better natural lang…

CCG SupertaggingChunking

Developing Politeness Annotated Corpus of Hindi Blogs

2014-05-01 · LREC 2014 5 · Ritesh Kumar

In this paper I discuss the creation and annotation of a corpus of Hindi blogs. The corpus consists of a total of over 479,000 blog posts and blog comments. It is annotated with the information about the politeness level…

How Are Spelling Errors Generated and Corrected? A Study of Corrected and Uncorrected Spelling Errors Using Keystroke Logs

2012-07-01 · ACL 2012 7 · Yukino Baba, Hisami Suzuki
Spelling CorrectionTransliteration

Mapping the Dialog Act Annotations of the LEGO Corpus into the Communicative Functions of ISO 24617-2

2016-12-05 · Eugénio Ribeiro, Ricardo Ribeiro, David Martins de Matos

In this paper we present strategies for mapping the dialog act annotations of the LEGO corpus into the communicative functions of the ISO 24617-2 standard. Using these strategies, we obtained an additional 347 dialogs an…

Automatic Recognition of Linguistic Replacements in Text Series Generated from Keystroke Logs

2016-05-01 · LREC 2016 5 · Daniel Couto-Vale, Stella Neumann, Paula Niemietz

This paper introduces a toolkit used for the purpose of detecting replacements of different grammatical and semantic structures in ongoing text production logged as a chronological series of computer interaction events (…