paper-with-me

홈 › Papers

\textit{Versteasch du mi?} Computational and Socio-Linguistic Perspectives on GenAI, LLMs, and Non-Standard Language

2026-03-30 · Verena Platzgummer, John McCrae, Sina Ahmadi arxiv

The design of Large Language Models and generative artificial intelligence has been shown to be "unfair" to less-spoken languages and to deepen the digital language divide. Critical sociolinguistic work has also argued that these technologies are not only made possible by prior socio-historical processes of linguistic standardisation, often grounded in European nationalist and colonial projects, but also exacerbate epistemologies of language as "monolithic, monolingual, syntactically standardized systems of meaning". In our paper, we draw on earlier work on the intersections of technology and language policy and bring our respective expertise in critical sociolinguistics and computational linguistics to bear on an interrogation of these arguments. We take two different complexes of non-standard linguistic varieties in our respective repertoires--South Tyrolean dialects, which are widely used in informal communication in South Tyrol, Italy, as well as varieties of Kurdish--as starting points to an interdisciplinary exploration of the intersections between GenAI and linguistic variation and standardisation. We discuss both how LLMs can be made to deal with nonstandard language from a technical perspective, and whether, when or how this can contribute to "democratic and decolonial digital and machine learning strategies", which has direct policy implications.

📄 PDF Abstract BibTeX arXiv:2603.28213

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inducing a lexicon of sociolinguistic variables from code-mixed text

2018-11-01 · WS 2018 11 · Philippa Shoemark, James Kirby, Sharon Goldwater

Sociolinguistics is often concerned with how variants of a linguistic item (e.g., \textit{nothing} vs. \textit{nothin{'}}) are used by different groups or in different situations. We introduce the task of inducing lexica…

Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation

2026-05-07 · Maximilian Maurer, Maximilian Linde, Gabriella Lapesa arxiv

Human label variation has been established as a central phenomenon in NLP: the perspectives different annotators have on the same item need to be embraced. Data collection practices thus shifted towards increasing the an…

Computational Sociolinguistics: A Survey

2015-08-30 · Dong Nguyen, A. Seza Doğruöz, Carolyn P. Rosé, Franciska de Jong

Language is a social phenomenon and variation is inherent to its social nature. Recently, there has been a surge of interest within the computational linguistics (CL) community in the social dimension of language. In thi…

Survey

Context-sensitive evaluation of automatic speech recognition: considering user experience & language variation

2021-04-01 · EACL (HCINLP) 2021 4 · Nina Markl, Catherine Lai

Commercial Automatic Speech Recognition (ASR) systems tend to show systemic predictive bias for marginalised speaker/user groups. We highlight the need for an interdisciplinary and context-sensitive approach to documenti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies

2023-01-05 · ACL 2021 5 · A. Seza Doğruöz, Sunayana Sitaram, Barbara E. Bullock, Almeida Jacqueline Toribio

The analysis of data in which multiple languages are represented has gained popularity among computational linguists in recent years. So far, much of this research focuses mainly on the improvement of computational metho…

Survey