paper-with-me

홈 › Papers

Modeling Orthographic Variation in Occitan's Dialects

2024-04-30 · Zachary William Hopton, Noëmi Aepli

Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a multilingual model with data from several Occitan dialects and conducted a series of experiments to assess the model's representations of these dialects. For evaluation purposes, we compiled a parallel lexicon encompassing four Occitan dialects. Intrinsic evaluations of the model's embeddings revealed that surface similarity between the dialects strengthened representations. When the model was further fine-tuned for part-of-speech tagging and Universal Dependency parsing, its performance was robust to dialectical variation, even when trained solely on part-of-speech data from a single dialect. Our findings suggest that large multilingual models minimize the need for spelling normalization during pre-processing.

📄 PDF Abstract BibTeX arXiv:2404.19315

Code (0)

등록된 구현이 없습니다.

Tasks

Dependency ParsingPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

A Four-Dialect Treebank for Occitan: Building Process and Parsing Experiments

2020-12-01 · VarDial (COLING) 2020 12 · Aleksandra Miletic, Myriam Bras, Marianne Vergez-Couret, Louise Esher 외

Occitan is a Romance language spoken mainly in the south of France. It has no official status in the country, it is not standardized and displays important diatopic variation resulting in a rich system of dialects. Recen…

A Computational Perspective on the Romanian Dialects

2016-05-01 · LREC 2016 5 · Alina Maria Ciobanu, Liviu P. Dinu

In this paper we conduct an initial study on the dialects of Romanian. We analyze the differences between Romanian and its dialects using the Swadesh list. We analyze the predictive power of the orthographic and phonetic…

Dialect IdentificationGeneral Classification

Automatic Pronunciation Generation by Utilizing a Semi-supervised Deep Neural Networks

2016-06-15 · Naoya Takahashi, Tofigh Naghibi, Beat Pfister

Phonemic or phonetic sub-word units are the most commonly used atomic elements to represent speech signals in modern ASRs. However they are not the optimal choice due to several reasons such as: large amount of effort re…

speech-recognitionSpeech Recognition

OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan

2022-10-01 · VarDial (COLING) 2022 10 · Aleksandra Miletic, Yves Scherrer

This paper presents OcWikiDisc, a new freely available corpus in Occitan, as well as language identification experiments on Occitan done as part of the corpus building process. Occitan is a regional language spoken mainl…

8kLanguage Identification

Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora

2025-09-22 · Robert Litschko, Verena Blaschke, Diana Burkhardt, Barbara Plank 외 arxiv

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address thi…