Inforex -- a web-based tool for text corpus management and semantic annotation
The aim of this paper is to present a system for semantic text annotation called Inforex. Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Named Entities (NE), anaphora, Word Sense Disambiguation (WSD) and relations between named entities. The system also supports manual text clean-up and automatic text pre-processing including text segmentation, morphosyntactic analysis and word selection for word sense annotation. Inforex can be accessed from any standard-compliant web browser supporting JavaScript. The user interface has a form of dynamic HTML pages using the AJAX technology. The server part of the system is written in PHP and the data is stored in MySQL database. The system make use of some external tools that are installed on the server or can be accessed via web services. The documents are stored in the database in the original format ― either plain text, XML or HTML. Tokenization and sentence segmentation is optional and is stored in a separate table. Tokens are stored as pairs of values representing indexes of first and last character of the tokens and sets of features representing the morpho-syntactic information.
Code (0)
등록된 구현이 없습니다.
Tasks
ManagementNamed Entity Recognition (NER)SentenceSentence segmentationtext annotationText SegmentationWord Sense DisambiguationSimilar Papers 제목 키워드 기반
Inforex --- a Collaborative Systemfor Text Corpora Annotation and Analysis Goes Open
In the paper we present the latest changes introduce to Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. One of the most important news is the release of source cod…
AttributeMorphological DisambiguationMorphological Taggingtext annotationInforex --- a collaborative system for text corpora annotation and analysis
We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…
Named Entity Recognition (NER)Word Sense DisambiguationLexiDB: Patterns \& Methods for Corpus Linguistic Database Management
LexiDB is a tool for storing, managing and querying corpus data. In contrast to other database management systems (DBMSs), it is designed specifically for text corpora. It improves on other corpus management systems (CMS…
ManagementRetrievalCooperating Tools for MWE Lexicon Management and Corpus Annotation
We present tools for lexicon and corpus management that offer cooperating functionality in corpus annotation. The former, named Cradle, stores a set of words and expressions where multi-word expressions are defined with …
ManagementPOSNLP Infrastructure for the Lithuanian Language
The Information System for Syntactic and Semantic Analysis of the Lithuanian language (lith. Lietuvi{\k{u}} kalbos sintaksin{\.e}s ir semantin{\.e}s analiz{\.e}s informacin{\.e} sistema, LKSSAIS) is the first infrastruct…
Managementnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1