paper-with-me

Papers

Creating a Lexicon of Bavarian Dialect by Means of Facebook Language Data and Crowdsourcing

2016-05-01 · LREC 2016 5 · Manuel Burghardt, Daniel Granvogl, Christian Wolff

Data acquisition in dialectology is typically a tedious task, as dialect samples of spoken language have to be collected via questionnaires or interviews. In this article, we suggest to use the {`}web as a corpus{''} approach for dialectology. We present a case study that demonstrates how authentic language data for the Bavarian dialect (ISO 639-3:bar) can be collected automatically from the social network Facebook. We also show that Facebook can be used effectively as a crowdsourcing platform, where users are willing to translate dialect words collaboratively in order to create a common lexicon of their Bavarian dialect. Key insights from the case study are summarized as {`}lessons learned{''}, together with suggestions for future enhancements of the lexicon creation approach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Using OntoLex-Lemon for Representing and Interlinking Lexicographic Collections of Bavarian Dialects

2020-05-01 · LREC 2020 5 · Yalemisew Abgaz

This paper describes the ongoing work in converting the lexicographic collection of a non-standard German language dataset (Bavarian Dialects) into a Linguistic Linked Open Data (LLOD) format. The collection is divided i…

Low-resource Bilingual Dialect Lexicon Induction with Large Language Models

2023-04-19 · Ekaterina Artemova, Barbara Plank

Bilingual word lexicons are crucial tools for multilingual natural language understanding and machine translation tasks, as they facilitate the mapping of words in one language to their synonyms in another language. To a…

Bilingual Lexicon InductionMachine TranslationNatural Language UnderstandingSemantic Similarity+3

Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora

2025-09-22 · Robert Litschko, Verena Blaschke, Diana Burkhardt, Barbara Plank 외 arxiv

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address thi…

Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data

2024-03-19 · Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm 외

Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with …

ArticlesDialect IdentificationDiversityMulti-Task Learning+4

Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study

2025-01-07 · Xaver Maria Krückl, Verena Blaschke, Barbara Plank

Reliable slot and intent detection (SID) is crucial in natural language understanding for applications like digital assistants. Encoder-only transformer models fine-tuned on high-resource languages generally perform well…

intent-classificationIntent ClassificationIntent DetectionLanguage Modelling+9