paper-with-me

Papers

Exploring Diachronic and Diatopic Changes in Dialect Continua: Tasks, Datasets and Challenges

2024-07-04 · Melis Çelikkol, Lydia Körber, Wei Zhao

Everlasting contact between language communities leads to constant changes in languages over time, and gives rise to language varieties and dialects. However, the communities speaking non-standard language are often overlooked by non-inclusive NLP technologies. Recently, there has been a surge of interest in studying diatopic and diachronic changes in dialect NLP, but there is currently no research exploring the intersection of both. Our work aims to fill this gap by systematically reviewing diachronic and diatopic papers from a unified perspective. In this work, we critically assess nine tasks and datasets across five dialects from three language families (Slavic, Romance, and Germanic) in both spoken and written modalities. The tasks covered are diverse, including corpus construction, dialect distance estimation, and dialect geolocation prediction, among others. Moreover, we outline five open challenges regarding changes in dialect use over time, the reliability of dialect datasets, the importance of speaker characteristics, limited coverage of dialects, and ethical considerations in data collection. We hope that our work sheds light on future research towards inclusive computational methods and datasets for language varieties and dialects.

📄 PDF Abstract BibTeX arXiv:2407.04010

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Corpora and Processing Tools for Non-standard Contemporary and Diachronic Balkan Slavic

2019-09-01 · RANLP 2019 9 · Teodora Vukovic, Nora Muheim, Olivier Winist{\"o}rfer, Ivan {\v{S}}imko 외

The paper describes three corpora of different varieties of BS that are currently being developed with the goal of providing data for the analysis of the diatopic and diachronic variation in non-standard Balkan Slavic. T…

LemmatizationPOS

Mapping Diatopic and Diachronic Variation in Spoken Czech: The ORTOFON and DIALEKT Corpora

2014-05-01 · LREC 2014 5 · Marie Kop{\v{r}}ivov{\'a}, Hana Gol{\'a}{\v{n}}ov{\'a}, Petra Klime{\v{s}}ov{\'a}, David Luke{\v{s}}

ORTOFON and DIALEKT are two corpora of spoken Czech (recordings + transcripts) which are currently being built at the Institute of the Czech National Corpus. The first one (ORTOFON) continues the tradition of the CNC{'}s…

Speech Recognition

DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish

2022-06-01 · LREC 2022 6 · Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde

We provide a novel dataset – DiaWUG – with judgements on diatopic lexical semantic variation for six Spanish variants in Europe and Latin America. In contrast to most previous meaning-based resources and studies on seman…

JeSemE: A Website for Exploring Diachronic Changes in Word Meaning and Emotion

2018-07-11 · Johannes Hellrich, Sven Buechel, Udo Hahn

We here introduce a substantially extended version of JeSemE, an interactive website for visually exploring computationally derived time-variant information on word meanings and lexical emotions assembled from five large…

Crowdsourcing Dialect Characterization through Twitter

2014-07-26 · Bruno Gonçalves, David Sánchez

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a care…