paper-with-me

Papers

Enhancing NER Performance in Low-Resource Pakistani Languages using Cross-Lingual Data Augmentation

2025-04-07 · Toqeer Ehsan, Thamar Solorio

Named Entity Recognition (NER), a fundamental task in Natural Language Processing (NLP), has shown significant advancements for high-resource languages. However, due to a lack of annotated datasets and limited representation in Pre-trained Language Models (PLMs), it remains understudied and challenging for low-resource languages. To address these challenges, we propose a data augmentation technique that generates culturally plausible sentences and experiments on four low-resource Pakistani languages; Urdu, Shahmukhi, Sindhi, and Pashto. By fine-tuning multilingual masked Large Language Models (LLMs), our approach demonstrates significant improvements in NER performance for Shahmukhi and Pashto. We further explore the capability of generative LLMs for NER and data augmentation using few-shot learning.

📄 PDF Abstract BibTeX arXiv:2504.08792

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationFew-Shot Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

Hurshch A Pakistani Choreographer And Director

2022-12-13 · 14-12 2022 12 · Hurshch, Rana Tahir Faisal

Hurshch is a Pakistani choreographer, model groomer, and director known for being the first model to work with the world's largest fashion company, FMD (The Fashion Model Directory).

BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment

2024-11-25 · Shaolei Zhang, Kehao Zhang, Qingkai Fang, Shoutao Guo 외

Large language models (LLMs), with their powerful generative capabilities and vast knowledge, empower various tasks in everyday life. However, these abilities are primarily concentrated in high-resource languages, leavin…

Language ModelingLanguage ModellingLarge Language ModelTransfer Learning

CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages

2025-09-21 · Wenhao Zhuang, Yuan Sun arxiv

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich lang…

Cross-Lingual TransferMachine Translation

Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

2024-02-19 · Yuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang 외

While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages. One common approach to mitigate this issu…

Transfer Learning

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

2024-04-03 · Zhongtao Miao, Qiyu Wu, Kaiyan Zhao, Zilong Wu 외

The field of cross-lingual sentence embeddings has recently experienced significant advancements, but research concerning low-resource languages has lagged due to the scarcity of parallel corpora. This paper shows that c…

RetrievalSentenceSentence EmbeddingSentence-Embedding+4