paper-with-me

Papers

Inforex -- a web-based tool for text corpus management and semantic annotation

2012-05-01 · LREC 2012 5 · Micha{\l} Marci{\'n}czuk, Jan Koco{\'n}, Bartosz Broda

The aim of this paper is to present a system for semantic text annotation called Inforex. Inforex is a web-based system designed for managing and annotating text corpora on the semantic level including annotation of Named Entities (NE), anaphora, Word Sense Disambiguation (WSD) and relations between named entities. The system also supports manual text clean-up and automatic text pre-processing including text segmentation, morphosyntactic analysis and word selection for word sense annotation. Inforex can be accessed from any standard-compliant web browser supporting JavaScript. The user interface has a form of dynamic HTML pages using the AJAX technology. The server part of the system is written in PHP and the data is stored in MySQL database. The system make use of some external tools that are installed on the server or can be accessed via web services. The documents are stored in the database in the original format ― either plain text, XML or HTML. Tokenization and sentence segmentation is optional and is stored in a separate table. Tokens are stored as pairs of values representing indexes of first and last character of the tokens and sets of features representing the morpho-syntactic information.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementNamed Entity Recognition (NER)SentenceSentence segmentationtext annotationText SegmentationWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Inforex --- a Collaborative Systemfor Text Corpora Annotation and Analysis Goes Open

2019-09-01 · RANLP 2019 9 · Micha{\l} Marci{\'n}czuk, Marcin Oleksy

In the paper we present the latest changes introduce to Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. One of the most important news is the release of source cod…

AttributeMorphological DisambiguationMorphological Taggingtext annotation

Inforex --- a collaborative system for text corpora annotation and analysis

2017-09-01 · RANLP 2017 9 · Micha{\l} Marci{\'n}czuk, Marcin Oleksy, Jan Koco{\'n}

We report a first major upgrade of Inforex {---} a web-based system for qualitative and collaborative text corpora annotation and analysis. Inforex is a part of Polish CLARIN infrastructure. It is integrated with a digit…

Named Entity Recognition (NER)Word Sense Disambiguation

LexiDB: Patterns \& Methods for Corpus Linguistic Database Management

2020-05-01 · LREC 2020 5 · Matthew Coole, Paul Rayson, John Mariani

LexiDB is a tool for storing, managing and querying corpus data. In contrast to other database management systems (DBMSs), it is designed specifically for text corpora. It improves on other corpus management systems (CMS…

ManagementRetrieval

Cooperating Tools for MWE Lexicon Management and Corpus Annotation

2018-08-01 · COLING 2018 8 · Yuji Matsumoto, Akihiko Kato, Hiroyuki Shindo, Toshio Morita

We present tools for lexicon and corpus management that offer cooperating functionality in corpus annotation. The former, named Cradle, stores a set of words and expressions where multi-word expressions are defined with …

ManagementPOS

NLP Infrastructure for the Lithuanian Language

2016-05-01 · LREC 2016 5 · Daiva Vitkut{\.e}-Ad{\v{z}}gauskien{\.e}, Andrius Utka, Darius Amilevi{\v{c}}ius, Tomas Krilavi{\v{c}}ius

The Information System for Syntactic and Semantic Analysis of the Lithuanian language (lith. Lietuvi{\k{u}} kalbos sintaksin{\.e}s ir semantin{\.e}s analiz{\.e}s informacin{\.e} sistema, LKSSAIS) is the first infrastruct…

Managementnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1