paper-with-me

Papers

Creation of a bottom-up corpus-based ontology for Italian Linguistics

2012-05-01 · LREC 2012 5 · Elisa Bianchi, Mirko Tavosanis, Emiliano Giovannetti

This paper describes the steps of construction of a shallow lexical ontology of Italian Linguistics, set to be used by a meta-search engine for query refinement. The ontology was constructed with the software Prot{\'e}g{\'e} 4.0.2 and is in OWL format; its construction has been carried out following the steps described in the well-known Ontology Learning From Text (OLFT) layer cake. The starting point was the automatic term extraction from a corpus of web documents concerning the domain of interest (304,000 words); as regards corpus construction, we describe the main criteria of the web documents selection and its critical points, concerning the definition of user profile and of degrees of specialisation. We describe then the process of term validation and construction of a glossary of terms of Italian Linguistics; afterwards, we outline the identification of synonymic chains and the main criteria of ontology design: top classes of ontology are Concept (containing taxonomy of concepts) and Terms (containing terms of the glossary as instances), while concepts are linked through part-whole and involved-role relation, both borrowed from Wordnet. Finally, we show some examples of the application of the ontology for query refinement.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Term Extraction

Similar Papers 제목 키워드 기반

CItA: an L1 Italian Learners Corpus to Study the Development of Writing Competence

2016-05-01 · LREC 2016 5 · Alessia Barbagli, Pietro Lucisano, Felice Dell{'}Orletta, Simonetta Montemagni 외

In this paper, we present the CItA corpus (Corpus Italiano di Apprendenti L1), a collection of essays written by Italian L1 learners collected during the first and second year of lower secondary school. The corpus was bu…

The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices

2024-11-10 · Ludovica Pannitto, Caterina Mauri

The paper presents an overview of initial design choices discussed towards the creation of a treebank for the Italian KIParla corpus

Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus

2025-09-23 · Chiara Alzetta, Serena Auriemma, Alessandro Bondielli, Luca Dini 외 arxiv

Over the past decade, Computational Linguistics (CL) and Natural Language Processing (NLP) have evolved rapidly, especially with the advent of Transformer-based Large Language Models (LLMs). This shift has transformed re…

Language Modelling

Towards the Creation of a Diachronic Corpus for Italian: A Case Study on the GDLI Quotations

2022-06-01 · LT4HALA (LREC) 2022 6 · Manuel Favaro, Elisa Guadagnini, Eva Sassolini, Marco Biffi 외

In this paper we describe some experiments related to a corpus derived from an authoritative historical Italian dictionary, namely the Grande dizionario della lingua italiana (‘Great Dictionary of Italian Language’, in s…

LemmatizationPOSPOS Tagging

A Large Interlinked Knowledge Graph of the Italian Cultural Heritage

2022-06-01 · LREC 2022 6 · Stefano Faralli, Andrea Lenzi, Paola Velardi

Knowledge is the lifeblood for a plethora of applications such as search, recommender systems and natural language understanding. Thanks to the efforts in the fields of Semantic Web and Linked Open Data a growing number …

Natural Language UnderstandingRecommendation Systems