Inforex --- a Collaborative Systemfor Text Corpora Annotation and Analysis Goes Open
In the paper we present the latest changes introduce to Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. One of the most important news is the release of source codes. Now the system is available on the GitHub repository (https://github.com/CLARIN-PL/Inforex) as an open source project. The system can be easily setup and run in a Docker container what simplifies the installation process. The major improvements include: semi-automatic text annotation, multilingual text preprocessing using CLARIN-PL web services, morphological tagging of XML documents, improved editor for annotation attribute, batch annotation attribute editor, morphological disambiguation, extended word sense annotation. This paper contains a brief description of the mentioned improvements. We also present two use cases in which various Inforex features were used and tested in real-life projects.
Code (1)
Tasks
AttributeMorphological DisambiguationMorphological Taggingtext annotationSimilar Papers 제목 키워드 기반
Inforex --- a collaborative system for text corpora annotation and analysis
We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…
Named Entity Recognition (NER)Word Sense DisambiguationInforex -- a web-based tool for text corpus management and semantic annotation
The aim of this paper is to present a system for semantic text annotation called Inforex. Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Name…
ManagementNamed Entity Recognition (NER)SentenceSentence segmentation+3TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts
The annotation of texts and other material in the field of digital humanities and Natural Language Processing (NLP) is a common task of research projects. At the same time, the annotation of corpora is certainly the most…
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
We introduce Co-DETECT (Collaborative Discovery of Edge cases in TExt ClassificaTion), a novel mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language mode…
Text ClassificationMultilingual corpora with coreferential annotation of person entities
This paper presents three corpora with coreferential annotation of person entities for Portuguese, Galician and Spanish. They contain coreference links between several types of pronouns (including elliptical, possessive,…
coreference-resolutionCoreference ResolutionOpen Information ExtractionRelation Extraction+1