paper-with-me

Papers

LAF-Fabric: a data analysis tool for Linguistic Annotation Framework with an application to the Hebrew Bible

2014-10-01 · Dirk Roorda, Gino Kalkman, Martijn Naaijer, Andreas van Cranenburgh

The Linguistic Annotation Framework (LAF) provides a general, extensible stand-off markup system for corpora. This paper discusses LAF-Fabric, a new tool to analyse LAF resources in general with an extension to process the Hebrew Bible in particular. We first walk through the history of the Hebrew Bible as text database in decennium-wide steps. Then we describe how LAF-Fabric may serve as an analysis tool for this corpus. Finally, we describe three analytic projects/workflows that benefit from the new LAF representation: 1) the study of linguistic variation: extract cooccurrence data of common nouns between the books of the Bible (Martijn Naaijer); 2) the study of the grammar of Hebrew poetry in the Psalms: extract clause typology (Gino Kalkman); 3) construction of a parser of classical Hebrew by Data Oriented Parsing: generate tree structures from the database (Andreas van Cranenburgh).

📄 PDF Abstract BibTeX arXiv:1410.0286

Code (1)

Dans-labs/text-fabric tf

Similar Papers 제목 키워드 기반

corpus-tools.org: An Interoperable Generic Software Tool Set for Multi-layer Linguistic Corpora

2016-05-01 · LREC 2016 5 · Stephan Druskat, Volker Gast, Thomas Krause, Florian Zipser

This paper introduces an open source, interoperable generic software tool set catering for the entire workflow of creation, migration, annotation, query and analysis of multi-layer linguistic corpora. It consists of four…

Standardisation and Interoperation of Morphosyntactic and Syntactic Annotation Tools for Spanish and their Annotations

2014-05-01 · LREC 2014 5 · Antonio Pareja-Lora, Guillermo C{\'a}rcamo-Escorza, Alicia Ballesteros-Calvo

Linguistic annotation tools and linguistic annotations are scarcely syntactically and/or semantically interoperable. Their low interoperability usually results from the number of factors taken into account in their devel…

Towards a Unified Tool for the Management of Data and Technologies in Field Linguistics and Computational Linguistics - LiFE

2022-06-01 · EURALI (LREC) 2022 6 · Siddharth Singh, Ritesh Kumar, Shyam Ratan, Sonal Sinha

The paper presents a new software - Linguistic Field Data Management and Analysis System - LiFE for endangered and low-resourced languages - an open-source, web-based linguistic data analysis and management application a…

Management

Iula2Standoff: a tool for creating standoff documents for the IULACT

2012-05-01 · LREC 2012 5 · Carlos Morell, Jorge Vivaldi, N{\'u}ria Bel

Due to the increase in the number and depth of analyses required over the text, like entity recognition, POS tagging, syntactic analysis, etc. the annotation in-line has become unpractical. In Natural Language Processing…

LemmatizationPOSPOS Tagging

Building Tamil Treebanks

2024-09-23 · Kengatharaiyer Sarveswaran

Treebanks are important linguistic resources, which are structured and annotated corpora with rich linguistic annotations. These resources are used in Natural Language Processing (NLP) applications, supporting linguistic…