paper-with-me

홈 › Papers

Producing Corpora of Medieval and Premodern Occitan

2019-04-26 · Jean-Baptiste Camps, Gilles Guilhem Couffignal

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a corpus of study tailored for their research, to enrich their data and to exploit them.Open optical character recognition (OCR) tools can be adapted to old prints, incunabula or even manuscripts, with usable results, allowing the rapid creation of textual corpora. The alternation of training and correction phases makes it possible to improve the quality of the results by rapidly accumulating raw text data. These can then be structured, for example in XML/TEI, and enriched.The enrichment of the texts with graphic or linguistic annotations can also be automated. These processes, known to linguists and functional for modern languages, present difficulties for languages such as Medieval Occitan, due in part to the absence of big enough lemmatized corpora. Suggestions for the creation of tools adapted to the considerable spelling variation of ancient languages will be presented, as well as experiments for the lemmatization of Medieval and Premodern Occitan.These techniques open the way for many exploitations. The much desired increase in the amount of available quality texts and data makes it possible to improve digital philology methods, if everyone takes the trouble to make their data freely available online and reusable.By exposing different technical solutions and some micro-analyses as examples, this paper aims to show part of what digital philology can offer to researchers in the Occitan domain, while recalling the ethical issues on which such practices are based.

📄 PDF Abstract BibTeX arXiv:1904.11815

Code (0)

등록된 구현이 없습니다.

Tasks

LemmatizationOptical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages

2026-05-09 · Matthias Schöffel, Esteban Garces Arias arxiv

Part-of-speech (POS) tagging for Medieval Romance languages remains challenging due to orthographic variation, morphological complexity, and limited annotated resources. This paper presents a systematic empirical evaluat…

Cross-Lingual TransferPOS Tagging

OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan

2022-10-01 · VarDial (COLING) 2022 10 · Aleksandra Miletic, Yves Scherrer

This paper presents OcWikiDisc, a new freely available corpus in Occitan, as well as language identification experiments on Occitan done as part of the corpus building process. Occitan is a regional language spoken mainl…

8kLanguage Identification

A Four-Dialect Treebank for Occitan: Building Process and Parsing Experiments

2020-12-01 · VarDial (COLING) 2020 12 · Aleksandra Miletic, Myriam Bras, Marianne Vergez-Couret, Louise Esher 외

Occitan is a Romance language spoken mainly in the south of France. It has no official status in the country, it is not standardized and displays important diatopic variation resulting in a rich system of dialects. Recen…

Building a treebank for Occitan: what use for Romance UD corpora?

2019-08-01 · WS 2019 8 · Aleks Miletic, ra, Myriam Bras, Louise Esher 외

Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard

2018-05-01 · LREC 2018 5 · Delphine Bernhard, Anne-Laure Ligozat, Fanny Martin, Myriam Bras 외