paper-with-me

홈 › Papers

Using full text indices for querying spoken language data

2020-05-01 · LREC 2020 5 · Elena Frick, Thomas Schmidt

As a part of the ZuMult-project, we are currently modelling a backend architecture that should provide query access to corpora from the Archive of Spoken German (AGD) at the Leibniz-Institute for the German Language (IDS). We are exploring how to reuse existing search engine frameworks providing full text indices and allowing to query corpora by one of the corpus query languages (QLs) established and actively used in the corpus research community. For this purpose, we tested MTAS - an open source Lucene-based search engine for querying on text with multilevel annotations. We applied MTAS on three oral corpora stored in the TEI-based ISO standard for transcriptions of spoken language (ISO 24624:2016). These corpora differ from the corpus data that MTAS was developed for, because they include interactions with two and more speakers and are enriched, inter alia, with timeline-based annotations. In this contribution, we report our test results and address issues that arise when search frameworks originally developed for querying written corpora are being transferred into the field of spoken language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Querying Interaction Structure: Approaches to Overlap in Spoken Language Corpora

2022-06-01 · LREC 2022 6 · Elena Frick, Thomas Schmidt, Henrike Helmer

In this paper, we address two problems in indexing and querying spoken language corpora with overlapping speaker contributions. First, we look into how token distance and token precedence can be measured when multiple pr…

BWT construction and search at the terabase scale

2024-09-01 · Heng Li

Motivation: Burrows-Wheeler Transform (BWT) is a common component in full-text indices. Initially developed for data compression, it is particularly powerful for encoding redundant sequences such as pangenome data. Howev…

Data Compression

Online Pandora's Box for Contextual LLM Cascading

2026-06-05 · Alexandre Belloni, Yan Chen, Yehua Wei arxiv

Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. In each period, a decision-maker observes a request context and faces…

Example-Based Treebank Querying with GrETEL--Now Also for Spoken Dutch

2013-05-01 · WS 2013 5 · Liesbeth Augustinus, V, Vincent eghinste, Ineke Schuurman 외

Scalable Semantic Querying of Text

2018-05-03 · Xiaolan Wang, Aaron Feng, Behzad Golshan, Alon Halevy 외

We present the KOKO system that takes declarative information extraction to a new level by incorporating advances in natural language processing techniques in its extraction language. KOKO is novel in that its extraction…

Articles