Modeling Orthographic Variation in Occitan's Dialects
Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a multilingual model with data from several Occitan dialects and conducted a series of experiments to assess the model's representations of these dialects. For evaluation purposes, we compiled a parallel lexicon encompassing four Occitan dialects. Intrinsic evaluations of the model's embeddings revealed that surface similarity between the dialects strengthened representations. When the model was further fine-tuned for part-of-speech tagging and Universal Dependency parsing, its performance was robust to dialectical variation, even when trained solely on part-of-speech data from a single dialect. Our findings suggest that large multilingual models minimize the need for spelling normalization during pre-processing.
Code (0)
등록된 구현이 없습니다.
Tasks
Dependency ParsingPart-Of-Speech TaggingSimilar Papers 제목 키워드 기반
A Four-Dialect Treebank for Occitan: Building Process and Parsing Experiments
Occitan is a Romance language spoken mainly in the south of France. It has no official status in the country, it is not standardized and displays important diatopic variation resulting in a rich system of dialects. Recen…
A Computational Perspective on the Romanian Dialects
In this paper we conduct an initial study on the dialects of Romanian. We analyze the differences between Romanian and its dialects using the Swadesh list. We analyze the predictive power of the orthographic and phonetic…
Dialect IdentificationGeneral ClassificationAutomatic Pronunciation Generation by Utilizing a Semi-supervised Deep Neural Networks
Phonemic or phonetic sub-word units are the most commonly used atomic elements to represent speech signals in modern ASRs. However they are not the optimal choice due to several reasons such as: large amount of effort re…
speech-recognitionSpeech RecognitionOcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan
This paper presents OcWikiDisc, a new freely available corpus in Occitan, as well as language identification experiments on Occitan done as part of the corpus building process. Occitan is a regional language spoken mainl…
8kLanguage IdentificationMake Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address thi…