Efficient Dependency Graph Matching with the IMS Open Corpus Workbench
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expressive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.
Code (0)
등록된 구현이 없습니다.
Tasks
Dependency ParsingGraph MatchingInformation RetrievalSentenceSimilar Papers 제목 키워드 기반
A Tree Extension for CoNLL-RDF
The technological bridges between knowledge graphs and natural language processing are of utmost importance for the future development of language technology. CoNLL-RDF is a technology that provides such a bridge for pop…
Knowledge GraphsSentencePaCMan : Parallel Corpus Management Workbench
The eIdentity Text Exploration Workbench
We work on tools to explore text contents and metadata of newspaper articles as provided by news archives. Our tool components are being integrated into an {``}Exploration Workbench{''} for Digital Humanities researchers…
ArticlesInformation RetrievalNamed Entity Recognition (NER)RetrievalSvarna: An Open Corpus Workbench for Modern Greek
This paper introduces Svarna, a free, open-source, web-based corpus workbench for modern Greek. Svarna integrates five databases covering various registers, institutional, literary, dialectal, social media, and historica…
Hypernym-LIBre: A Free Web-based Corpus for Hypernym Detection
In this paper, we describe a new web-based corpus for hypernym detection. It consists of 32 GB of high quality english paragraphs along with their part-of-speech tagged and dependency parsed versions. For hypernym detect…
POS