paper-with-me

Papers

Standardizing linguistic data: method and tools for annotating (pre-orthographic) French

2020-11-22 · Simon Gabay, Thibault Clérice, Jean-Baptiste Camps, Jean-Baptiste Tanguy, Matthias Gille-Levenson

With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the interoperability of the data produced, despite diachronic variations. In the present paper, we describe both methodologically (by proposing annotation principles) and technically (by creating the required training data and the relevant models) the production of a linguistic tagger for (early) modern French (16-18th c.), taking as much as possible into account already existing standards for contemporary and, especially, medieval French.

📄 PDF Abstract BibTeX arXiv:2011.11074

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

EXMARaLDA and the FOLK tools --- two toolsets for transcribing and annotating spoken language

2012-05-01 · LREC 2012 5 · Thomas Schmidt

This paper presents two toolsets for transcribing and annotating spoken language: the EXMARaLDA system, developed at the University of Hamburg, and the FOLK tools, developed at the Institute for the German Language in Ma…

Annotating and Learning Morphological Segmentation of Egyptian Colloquial Arabic

2012-05-01 · LREC 2012 5 · Emad Mohamed, Behrang Mohit, Kemal Oflazer

We present an annotation and morphological segmentation scheme for Egyptian Colloquial Arabic (ECA) in which we annotate user-generated content that significantly deviates from the orthographic and grammatical rules of M…

General ClassificationPart-Of-Speech Tagging

Speeding up corpus development for linguistic research: language documentation and acquisition in Romansh Tuatschin

2017-08-01 · WS 2017 8 · G{\'e}raldine Walther, Beno{\^\i}t Sagot

In this paper, we present ongoing work for developing language resources and basic NLP tools for an undocumented variety of Romansh, in the context of a language documentation and language acquisition project. Our tools …

Language AcquisitionSpelling Correction

Standardizing a Component Metadata Infrastructure

2012-05-01 · LREC 2012 5 · Daan Broeder, Dieter van Uytvanck, Maria Gavrilidou, Thorsten Trippel 외

This paper describes the status of the standardization efforts of a Component Metadata approach for describing Language Resources with metadata. Different linguistic and Language {\&} Technology communities as CLARIN, ME…

Turkish Resources for Visual Word Recognition

2014-05-01 · LREC 2014 5 · Beg{\"u}m Erten, Cem Bozsahin, Deniz Zeyrek

We report two tools to conduct psycholinguistic experiments on Turkish words. KelimetriK allows experimenters to choose words based on desired orthographic scores of word frequency, bigram and trigram frequency, ON, OLD2…

Language ModellingSpeech Recognition