paper-with-me

홈 › Papers

LexiDB: Patterns \& Methods for Corpus Linguistic Database Management

2020-05-01 · LREC 2020 5 · Matthew Coole, Paul Rayson, John Mariani

LexiDB is a tool for storing, managing and querying corpus data. In contrast to other database management systems (DBMSs), it is designed specifically for text corpora. It improves on other corpus management systems (CMSs) because data can be added and deleted from corpora on the fly with the ability to add live data to existing corpora. LexiDB sits between these two categories of DBMSs and CMSs, more specialised to language data than a general purpose DBMS but more flexible than a traditional static corpus management system. Previous work has demonstrated the scalability of LexiDB in response to the growing need to be able to scale out for ever growing corpus datasets. Here, we present the patterns and methods developed in LexiDB for storage, retrieval and querying of multi-level annotated corpus data. These techniques are evaluated and compared to an existing CMS (Corpus Workbench CWB - CQP) and indexer (Lucene). We find that LexiDB consistently outperforms existing tools for corpus queries. This is particularly apparent with large corpora and when handling queries with large result sets

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementRetrieval

Similar Papers 제목 키워드 기반

Infrastructure for Semantic Annotation in the Genomics Domain

2020-05-01 · LREC 2020 5 · Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani 외

We describe a novel super-infrastructure for biomedical text mining which incorporates an end-to-end pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature, co…

Retrieval

Unfinished Business: Construction and Maintenance of a Semantically Tagged Historical Parliamentary Corpus, UK Hansard from 1803 to the present day

2020-05-01 · LREC 2020 5 · Matthew Coole, Paul Rayson, John Mariani

Creating, curating and maintaining modern political corpora is becoming an ever more involved task. As interest from various social bodies and the general public in political discourse grows so too does the need to enric…

Part-Of-Speech Tagging

A Grammar-informed Corpus-based Sentence Database for Linguistic and Computational Studies

2012-05-01 · LREC 2012 5 · Hongzhi Xu, Helen Kai-yun Chen, Chu-Ren Huang, Qin Lu 외

We adopt the corpus-informed approach to example sentence selections for the construction of a reference grammar. In the process, a database containing sentences that are carefully selected by linguistic experts includin…

Chinese Word SegmentationPOSPOS TaggingSentence

The Slovene BNSI Broadcast News database and reference speech corpus GOS: Towards the uniform guidelines for future work

2014-05-01 · LREC 2014 5 · Andrej {\v{Z}}gank, Ana Zwitter Vitez, Darinka Verdonik

The aim of the paper is to search for common guidelines for the future development of speech databases for less resourced languages in order to make them the most useful for both main fields of their use, linguistic rese…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

The ACQDIV Corpus Database and Aggregation Pipeline

2020-05-01 · LREC 2020 5 · Anna Jancso, Steven Moran, Sabine Stoll

We present the ACQDIV corpus database and aggregation pipeline, a tool developed as part of the European Research Council (ERC) funded project ACQDIV, which aims to identify the universal cognitive processes that allow c…

Language Acquisition