paper-with-me

홈 › Papers

Connecting Ideas in 'Lower-Resource' Scenarios: NLP for National Varieties, Creoles and Other Low-resource Scenarios

2024-09-19 · Aditya Joshi, Diptesh Kanojia, Heather Lent, Hour Kaing, Haiyue Song

Despite excellent results on benchmarks over a small subset of languages, large language models struggle to process text from languages situated in lower-resource' scenarios such as dialects/sociolects (national or social varieties of a language), Creoles (languages arising from linguistic contact between multiple languages) and other low-resource languages. This introductory tutorial will identify common challenges, approaches, and themes in natural language processing (NLP) research for confronting and overcoming the obstacles inherent to data-poor contexts. By connecting past ideas to the present field, this tutorial aims to ignite collaboration and cross-pollination between researchers working in these scenarios. Our notion of lower-resource' broadly denotes the outstanding lack of data required for model training - and may be applied to scenarios apart from the three covered in the tutorial.

📄 PDF Abstract BibTeX arXiv:2409.12683

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ELRI: A Decentralised Network of National Relay Stations to Collect, Prepare and Share Language Resources

2020-05-01 · LREC 2020 5 · Thierry Etchegoyhen, Borja Anza Porras, Andoni Azpeitia, Eva Mart{\'\i}nez Garcia 외

We describe the European Language Resource Infrastructure (ELRI), a decentralised network to help collect, prepare and share language resources. The infrastructure was developed within a project co-funded by the Connecti…

Translation

The Role of Computing Resources in Publishing Foundation Model Research

2025-10-15 · Yuexing Hao, Yue Huang, Haoran Zhang, Chenyang Zhao 외 arxiv

Cutting-edge research in Artificial Intelligence (AI) requires considerable resources, including Graphics Processing Units (GPUs), data, and human resources. In this paper, we evaluate of the relationship between these r…

The Polish Sejm Corpus

2012-05-01 · LREC 2012 5 · Maciej Ogrodniczuk

This document presents the first edition of the Polish Sejm Corpus -- a new specialized resource containing transcribed, automatically annotated utterances of the Members of Polish Sejm (lower chamber of the Polish Parli…

SentenceWord Sense Disambiguation

Multi-Sentence Argument Linking

2019-11-09 · ACL 2020 6 · Seth Ebner, Patrick Xia, Ryan Culkin, Kyle Rawlins 외

We present a novel document-level model for finding argument spans that fill an event's roles, connecting related ideas in sentence-level semantic role labeling and coreference resolution. Because existing datasets for c…

coreference-resolutionCoreference ResolutionSemantic Role LabelingSentence+1

Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition

2023-05-19 · Siyuan Feng, Ming Tu, Rui Xia, Chuanzeng Huang 외

We improve low-resource ASR by integrating the ideas of multilingual training and self-supervised learning. Concretely, we leverage an International Phonetic Alphabet (IPA) multilingual model to create frame-level pseudo…

DiversitySelf-Supervised Learningspeech-recognitionSpeech Recognition