paper-with-me

Papers

Modelling and Annotating Interlinear Glossed Text from 280 Different Endangered Languages as Linked Data with LIGT

2020-12-01 · COLING (LAW) 2020 12 · Sebastian Nordhoff

This paper reports on the harvesting, analysis, and enrichment of 20k documents from 4 different endangered language archives in 300 different low-resource languages. The documents are heterogeneous as to their provenance (holding archive, language, geographical area, creator) and internal structure (annotation types, metalanguages), but they have the ELAN-XML format in common. Typical annotations include sentence-level translations, morpheme-segmentation, morpheme-level translations, and parts-of-speech. The ELAN-format gives a lot of freedom to document creators, and hence the data set is very heterogeneous. We use regularities in the ELAN format to arrive at a common internal representation of sentences, words, and morphemes, with translations into one or more additional languages. Building upon the paradigm of Linguistic Linked Open Data (LLOD, Chiarcos, Nordhoff, et al. 2012), the document elements receive unique identifiers and are linked to other resources such as Glottolog for languages, Wikidata for semantic concepts, and the Leipzig Glossing Rules list for category abbreviations. We provide an RDF export in the LIGT format (Chiarcos & Ionov 2019), enabling uniform and interoperable access with some semantic enrichments to a formerly disparate resource type difficult to access. Two use cases (semantic search and colexification) are presented to show the viability of the approach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Extracting Interlinear Glossed Text from LaTeX Documents

2016-05-01 · LREC 2016 5 · Mathias Schenner, Sebastian Nordhoff

We present texigt, a command-line tool for the extraction of structured linguistic data from LaTeX source documents, and a language resource that has been generated using this tool: a corpus of interlinear glossed text (…

IMTVault: Extracting and Enriching Low-resource Language Interlinear Glossed Text from Grammatical Descriptions and Typological Survey Articles

2022-06-01 · LDL (ACL) 2022 6 · Sebastian Nordhoff, Thomas Krämer

Many NLP resources and programs focus on a handful of major languages. But there are thousands of languages with low or no resources available as structured data. This paper shows the extraction of 40k examples with inte…

ArticlesTranslation

Automating Gloss Generation in Interlinear Glossed Text

2020-01-01 · SCiL 2020 1 · Angelina McMillan-Major

Improving Dependency Parsing with Interlinear Glossed Text and Syntactic Projection

2012-12-01 · COLING 2012 12 · Ryan Georgi, Fei Xia, William Lewis
Dependency ParsingWord Alignment

Enhanced and Portable Dependency Projection Algorithms Using Interlinear Glossed Text

2013-08-01 · ACL 2013 8 · Ryan Georgi, Fei Xia, William D. Lewis
Word Alignment