Standardizing linguistic data: method and tools for annotating (pre-orthographic) French
With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the interoperability of the data produced, despite diachronic variations. In the present paper, we describe both methodologically (by proposing annotation principles) and technically (by creating the required training data and the relevant models) the production of a linguistic tagger for (early) modern French (16-18th c.), taking as much as possible into account already existing standards for contemporary and, especially, medieval French.
Code (0)
등록된 구현이 없습니다.
Tasks
POSSimilar Papers 제목 키워드 기반
EXMARaLDA and the FOLK tools --- two toolsets for transcribing and annotating spoken language
This paper presents two toolsets for transcribing and annotating spoken language: the EXMARaLDA system, developed at the University of Hamburg, and the FOLK tools, developed at the Institute for the German Language in Ma…
Annotating and Learning Morphological Segmentation of Egyptian Colloquial Arabic
We present an annotation and morphological segmentation scheme for Egyptian Colloquial Arabic (ECA) in which we annotate user-generated content that significantly deviates from the orthographic and grammatical rules of M…
General ClassificationPart-Of-Speech TaggingSpeeding up corpus development for linguistic research: language documentation and acquisition in Romansh Tuatschin
In this paper, we present ongoing work for developing language resources and basic NLP tools for an undocumented variety of Romansh, in the context of a language documentation and language acquisition project. Our tools …
Language AcquisitionSpelling CorrectionStandardizing a Component Metadata Infrastructure
This paper describes the status of the standardization efforts of a Component Metadata approach for describing Language Resources with metadata. Different linguistic and Language {\&} Technology communities as CLARIN, ME…
Turkish Resources for Visual Word Recognition
We report two tools to conduct psycholinguistic experiments on Turkish words. KelimetriK allows experimenters to choose words based on desired orthographic scores of word frequency, bigram and trigram frequency, ON, OLD2…
Language ModellingSpeech Recognition