Creating a Lexicon of Bavarian Dialect by Means of Facebook Language Data and Crowdsourcing
Data acquisition in dialectology is typically a tedious task, as dialect samples of spoken language have to be collected via questionnaires or interviews. In this article, we suggest to use the {`}web as a corpus{''} approach for dialectology. We present a case study that demonstrates how authentic language data for the Bavarian dialect (ISO 639-3:bar) can be collected automatically from the social network Facebook. We also show that Facebook can be used effectively as a crowdsourcing platform, where users are willing to translate dialect words collaboratively in order to create a common lexicon of their Bavarian dialect. Key insights from the case study are summarized as {`}lessons learned{''}, together with suggestions for future enhancements of the lexicon creation approach.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Using OntoLex-Lemon for Representing and Interlinking Lexicographic Collections of Bavarian Dialects
This paper describes the ongoing work in converting the lexicographic collection of a non-standard German language dataset (Bavarian Dialects) into a Linguistic Linked Open Data (LLOD) format. The collection is divided i…
Low-resource Bilingual Dialect Lexicon Induction with Large Language Models
Bilingual word lexicons are crucial tools for multilingual natural language understanding and machine translation tasks, as they facilitate the mapping of words in one language to their synonyms in another language. To a…
Bilingual Lexicon InductionMachine TranslationNatural Language UnderstandingSemantic Similarity+3Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address thi…
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with …
ArticlesDialect IdentificationDiversityMulti-Task Learning+4Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study
Reliable slot and intent detection (SID) is crucial in natural language understanding for applications like digital assistants. Encoder-only transformer models fine-tuned on high-resource languages generally perform well…
intent-classificationIntent ClassificationIntent DetectionLanguage Modelling+9