paper-with-me

Papers

``Voices of the Great War'': A Richly Annotated Corpus of Italian Texts on the First World War

2020-05-01 · LREC 2020 5 · Federico Boschetti, Irene De Felice, Stefano Dei Rossi, Felice Dell{'}Orletta, Michele Di Giorgio, Martina Miliani, Lucia C. Passaro, Angelica Puddu, Giulia Venturi, Nicola Labanca, Aless Lenci, ro, Simonetta Montemagni

{`}Voices of the Great War{''} is the first large corpus of Italian historical texts dating back to the period of First World War. This corpus differs from other existing resources in several respects. First, from the linguistic point of view it gives account of the wide range of varieties in which Italian was articulated in that period, namely from a diastratic (educated vs. uneducated writers), diaphasic (low/informal vs. high/formal registers) and diatopic (regional varieties, dialects) points of view. From the historical perspective, through a collection of texts belonging to different genres it represents different views on the war and the various styles of narrating war events and experiences. The final corpus is balanced along various dimensions, corresponding to the textual genre, the language variety used, the author type and the typology of conveyed contents. The corpus is fully annotated with lemmas, part-of-speech, terminology, and named entities. Significant corpus samples representative of the different {`}voices{''} have also been enriched with meta-linguistic and syntactic information. The layer of syntactic annotation forms the first nucleus of an Italian historical treebank complying with the Universal Dependencies standard. The paper illustrates the final resource, the methodology and tools used to build it, and the Web Interface for navigating it.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

2026-06-10 · Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini arxiv

We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. The corpus, in its current version, is composed of approximately 4 million clini…

Information Extraction

Towards the Creation of a Diachronic Corpus for Italian: A Case Study on the GDLI Quotations

2022-06-01 · LT4HALA (LREC) 2022 6 · Manuel Favaro, Elisa Guadagnini, Eva Sassolini, Marco Biffi 외

In this paper we describe some experiments related to a corpus derived from an authoritative historical Italian dictionary, namely the Grande dizionario della lingua italiana (‘Great Dictionary of Italian Language’, in s…

LemmatizationPOSPOS Tagging

EMOVO Corpus: an Italian Emotional Speech Database

2014-05-01 · LREC 2014 5 · Giovanni Costantini, Iacopo Iaderola, Andrea Paoloni, Massimiliano Todisco

This article describes the first emotional corpus, named EMOVO, applicable to Italian language,. It is a database built from the voices of up to 6 actors who played 14 sentences simulating 6 emotional states (disgust, fe…

Emotion RecognitionSpeech Emotion Recognition

The making of the Litkey Corpus, a richly annotated longitudinal corpus of German texts written by primary school children

2019-08-01 · WS 2019 8 · Ronja Laarmann-Quante, Stefanie Dipper, Eva Belke

To date, corpus and computational linguistic work on written language acquisition has mostly dealt with second language learners who have usually already mastered orthography acquisition in their first language. In this …

Language AcquisitionPOS

EventNet-ITA: Italian Frame Parsing for Events

2023-05-18 · Marco Rovera

This paper introduces EventNet-ITA, a large, multi-domain corpus annotated full-text with event frames for Italian. Moreover, we present and thoroughly evaluate an efficient multi-label sequence labeling approach for Fra…