paper-with-me

홈 › Papers

Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies

2025-12-16 · Ekaterina Artemova, Laurie Burchell, Daryna Dementieva, Shu Okabe, Mariya Shmatova, Pedro Ortiz Suarez arxiv

This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful language technologies. Participants will walk away with a practical toolkit for building end-to-end NLP pipelines for underrepresented languages -- from data collection and web crawling to parallel sentence mining, machine translation, and downstream applications such as text classification and multimodal reasoning. The tutorial presents strategies for tackling the challenges of data scarcity and cultural variance, offering hands-on methods and modeling frameworks. We will focus on fair, reproducible, and community-informed development approaches, grounded in real-world scenarios. We will showcase a diverse set of use cases covering over 10 languages from different language families and geopolitical contexts, including both digitally resource-rich and severely underrepresented languages.

📄 PDF Abstract BibTeX arXiv:2512.14576

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningMachine TranslationText Classification

Similar Papers 제목 키워드 기반

MassiveSumm: a very large-scale, very multilingual, news summarisation dataset

2021-11-01 · EMNLP 2021 11 · Daniel Varab, Natalie Schluter

Current research in automatic summarisation is unapologetically anglo-centered–a persistent state-of-affairs, which also predates neural net approaches. High-quality automatic summarisation datasets are notoriously expen…

Articles

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions

2026-05-18 · Sicheng Jin, Dipankar Srirag, Aditya Joshi arxiv

While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to a…

Speech Recognition

Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo

2025-01-19 · Audrey Mbogho, Quin Awuor, Andrew Kipkebut, Lilian Wanzare 외

Natural Language Processing is a crucial frontier in artificial intelligence, with broad applications in many areas, including public health, agriculture, education, and commerce. However, due to the lack of substantial …

The Impact of Digitalisation and Sustainability on Inclusiveness: Inclusive Growth Determinants

2025-01-14 · Radu Rusu, Camelia Oprean-Stan

Inclusiveness and economic development have been slowed by the pandemics and military conflicts. This study investigates the main determinants of inclusiveness at the European level. A multi-method approach is used, with…

Management

Towards dialect-inclusive recognition in a low-resource language: are balanced corpora the answer?

2023-07-14 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl 외

ASR systems are generally built for the spoken 'standard', and their performance declines for non-standard dialects/varieties. This is a problem for a language like Irish, where there is no single spoken standard, but ra…

Diagnostic