Building a treebank for Occitan: what use for Romance UD corpora?
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Four-Dialect Treebank for Occitan: Building Process and Parsing Experiments
Occitan is a Romance language spoken mainly in the south of France. It has no official status in the country, it is not standardized and displays important diatopic variation resulting in a rich system of dialects. Recen…
Tools for Digital Humanities: Enabling Access to the Old Occitan Romance of Flamenca
Building a Universal Dependencies Treebank for Occitan
This paper outlines the ongoing effort of creating the first treebank for Occitan, a low-ressourced regional language spoken mainly in the south of France. We briefly present the global context of the project and report …
POSDo not neglect related languages: The case of low-resource Occitan cross-lingual word embeddings
Cross-lingual word embeddings (CLWEs) have proven indispensable for various natural language processing tasks, e.g., bilingual lexicon induction (BLI). However, the lack of data often impairs the quality of representatio…
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord EmbeddingsOcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan
This paper presents OcWikiDisc, a new freely available corpus in Occitan, as well as language identification experiments on Occitan done as part of the corpus building process. Occitan is a regional language spoken mainl…
8kLanguage Identification