E:Calm Resource: a Resource for Studying Texts Produced by French Pupils and Students
The E:Calm resource is constructed from French student texts produced in a variety of usual contexts of teaching. The distinction of the E:Calm resource is to provide an ecological data set that gives a broad overview of texts written at elementary school, high school and university. This paper describes the whole data processing: encoding of the main graphical aspects of the handwritten primary sources according to the TEI-P5 norm; spelling standardizing; POS tagging and syntactic parsing evaluation.
Code (0)
등록된 구현이 없습니다.
Tasks
POSPOS TaggingSimilar Papers 제목 키워드 기반
Extracting Linguistic Resources from the Web for Concept-to-Text Generation
Many concept-to-text generation systems require domain-specific linguistic resources to produce high quality texts, but manually constructing these resources can be tedious and costly. Focusing on NaturalOWL, a publicly …
Concept-To-Text GenerationSentenceText GenerationTheoretical and Methodological Framework for Studying Texts Produced by Large Language Models
This paper addresses the conceptual, methodological and technical challenges in studying large language models (LLMs) and the texts they produce from a quantitative linguistics perspective. It builds on a theoretical fra…
New language resources for the Pashto language
This paper reports on the development of new language resources for the Pashto language, a very low-resource language spoken in Afghanistan and Pakistan. In the scope of a multilingual data collection project, three larg…
Machine TranslationTranslationDimStance: Multilingual Datasets for Dimensional Stance Analysis
Stance detection is an established task that classifies an author's attitude toward a specific target into categories such as Favor, Neutral, and Against. Beyond categorical stance labels, we leverage a long-established …
Stance DetectionChiSCor: A Corpus of Freely Told Fantasy Stories by Dutch Children for Computational Linguistics and Cognitive Science
In this resource paper we release ChiSCor, a new corpus containing 619 fantasy stories, told freely by 442 Dutch children aged 4-12. ChiSCor was compiled for studying how children render character perspectives, and unrav…
LEMMAvalid