Evaluating Query Languages for a Corpus Processing System
This paper documents a pilot study conducted as part of the development of a new corpus processing system at the Institut f{\"u}r Deutsche Sprache in Mannheim and in the context of the ISO TC37 SC4/WG6 activity on the suggested work item proposal Corpus Query Lingua Franca. We describe the first phase of our research: the initial formulation of functionality criteria for query language evaluation and the results of the application of these criteria to three representatives of corpus query languages, namely COSMAS II, Poliqarp, and ANNIS QL. In contrast to previous works on query language evaluation that compare a range of existing query languages against a small number of queries, our approach analyses only three query languages against criteria derived from a suite of 300 use cases that cover diverse aspects of linguistic research.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Corpus Query Middleware of Tomorrow -- A Proposal for a Hybrid Corpus Query Architecture
Development of dozens of specialized corpus query systems and languages over the past decades has let to a diverse but also fragmented landscape. Today we are faced with a plethora of query tools that each provide unique…
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Semantic code search is the task of retrieving relevant code given a natural language query. While related to other information retrieval tasks, it requires bridging the gap between the language used in code (often abbre…
4kCode SearchInformation RetrievalNatural Language Queries+1Improving corpus search via parsing
In this paper, we describe an addition to the corpus query system Kontext that enables to enhance the search using syntactic attributes in addition to the existing features, mainly lemmas and morphological categories. We…
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying
Corpus query systems exist to address the multifarious information needs of any person interested in the content of annotated corpora. In this role they play an important part in making those resources usable for a wider…
A Parallel Corpus for Evaluating Machine Translation between Arabic and European Languages
We present Arab-Acquis, a large publicly available dataset for evaluating machine translation between 22 European languages and Arabic. Arab-Acquis consists of over 12,000 sentences from the JRC-Acquis (Acquis Communauta…
BenchmarkingMachine TranslationTranslation