paper-with-me

홈 › Papers

Exploring NLP Benchmarks in an Extremely Low-Resource Setting

2025-09-04 · Ulin Nuha, Adam Jatowt arxiv

The effectiveness of Large Language Models (LLMs) diminishes for extremely low-resource languages, such as indigenous languages, primarily due to the lack of labeled data. Despite growing interest, the availability of high-quality natural language processing (NLP) datasets for these languages remains limited, making it difficult to develop robust language technologies. This paper addresses such gap by focusing on Ladin, an endangered Romance language, specifically targeting the Val Badia variant. Leveraging a small set of parallel Ladin-Italian sentence pairs, we create synthetic datasets for sentiment analysis and multiple-choice question answering (MCQA) by translating monolingual Italian data. To ensure linguistic quality and reliability, we apply rigorous filtering and back-translation procedures in our method. We further demonstrate that incorporating these synthetic datasets into machine translation training leads to substantial improvements over existing Italian-Ladin translation baselines. Our contributions include the first publicly available sentiment analysis and MCQA datasets for Ladin, establishing foundational resources that can support broader NLP research and downstream applications for this underrepresented language.

📄 PDF Abstract BibTeX arXiv:2509.03962

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentiment AnalysisQuestion Answering

Similar Papers 제목 키워드 기반

What are the limits of cross-lingual dense passage retrieval for low-resource languages?

2024-08-21 · Jie Wu, Zhaochun Ren, Suzan Verberne

In this paper, we analyze the capabilities of the multi-lingual Dense Passage Retriever (mDPR) for extremely low-resource languages. In the Cross-lingual Open-Retrieval Answer Generation (CORA) pipeline, mDPR achieves su…

Answer GenerationLanguage ModelingLanguage ModellingPassage Retrieval+2

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource Languages

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Neural Machine Translation (NMT) models have been effective on large bilingual datasets. However, the existing methods and techniques show that the model's performance is highly dependent on the number of examples in tra…

Machine TranslationNMTTranslation

Dict-NMT: Bilingual Dictionary based NMT for Extremely Low Resource Languages

2022-06-09 · Nalin Kumar, Deepak Kumar, Subhankar Mishra

Neural Machine Translation (NMT) models have been effective on large bilingual datasets. However, the existing methods and techniques show that the model's performance is highly dependent on the number of examples in tra…

Machine TranslationNMTTranslation

Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages

2024-09-13 · Yao-Fei Cheng, Li-Wei Chen, Hung-Shin Lee, Hsin-Min Wang

This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual TransferData Augmentation+4

On Optimal Transformer Depth for Low-Resource Language Translation

2020-04-09 · Elan van Biljon, Arnu Pretorius, Julia Kreutzer

Transformers have shown great promise as an approach to Neural Machine Translation (NMT) for low-resource languages. However, at the same time, transformer models remain difficult to optimize and require careful tuning o…

Low Resource NMTMachine TranslationNMTTranslation