paper-with-me

Papers

A Survey of Corpora for Germanic Low-Resource Languages and Dialects

2023-04-19 · Verena Blaschke, Hinrich Schütze, Barbara Plank

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized low-resource languages. Even within branches of major language families, often considered well-researched, little is known about the extent and type of available resources and what the major NLP challenges are for these language varieties. The first step to address this situation is a systematic survey of available corpora (most importantly, annotated corpora, which are particularly valuable for NLP research). Focusing on Germanic low-resource language varieties, we provide such a survey in this paper. Except for geolocation (origin of speaker or document), we find that manually annotated linguistic resources are sparse and, if they exist, mostly cover morphosyntax. Despite this lack of resources, we observe that interest in this area is increasing: there is active development and a growing research community. To facilitate research, we make our overview of over 80 corpora publicly available. We share a companion website of this overview at https://github.com/mainlp/germanic-lrl-corpora .

📄 PDF Abstract BibTeX arXiv:2304.09805

Code (2)

mainlp/germanic-lrl-corpora 공식 구현
mainlp/wikistats 공식 구현

Tasks

Survey

Similar Papers 제목 키워드 기반

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

2026-04-20 · Raghvendra Kumar, Devankar Raj, Sriparna Saha arxiv

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates r…

Domain Generalization

Improving Zero-shot Cross-lingual Transfer between Closely Related Languages by injecting Character-level Noise

2021-09-14 · Findings (ACL) 2022 5 · Noëmi Aepli, Rico Sennrich

Cross-lingual transfer between a high-resource language and its dialects or closely related language varieties should be facilitated by their similarity. However, current approaches that operate in the embedding space do…

Cross-Lingual TransferPOSPOS TaggingZero-Shot Cross-Lingual Transfer

LuxBank: The First Universal Dependency Treebank for Luxembourgish

2024-11-07 · Alistair Plum, Caroline Döhmer, Emilia Milano, Anne-Marie Lutgen 외

The Universal Dependencies (UD) project has significantly expanded linguistic coverage across 161 languages, yet Luxembourgish, a West Germanic language spoken by approximately 400,000 people, has remained absent until n…

Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview

2020-10-14 · Alena Butryna, Shan-Hui Cathy Chu, Isin Demirsahin, Alexander Gutkin 외

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building tex…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Cyberbullying Detection for Low-resource Languages and Dialects: Review of the State of the Art

2023-08-30 · Tanjim Mahmud, Michal Ptaszynski, Juuso Eronen, Fumito Masui

The struggle of social media platforms to moderate content in a timely manner, encourages users to abuse such platforms to spread vulgar or abusive language, which, when performed repeatedly becomes cyberbullying a socia…

Abusive Language