paper-with-me

홈 › Papers

Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin

2024-04-28 · Pin-Jie Lin, Merel Scholman, Muhammed Saeed, Vera Demberg

Nigerian Pidgin is an English-derived contact language and is traditionally an oral language, spoken by approximately 100 million people. No orthographic standard has yet been adopted, and thus the few available Pidgin datasets that exist are characterised by noise in the form of orthographic variations. This contributes to under-performance of models in critical NLP tasks. The current work is the first to describe various types of orthographic variations commonly found in Nigerian Pidgin texts, and model this orthographic variation. The variations identified in the dataset form the basis of a phonetic-theoretic framework for word editing, which is used to generate orthographic variations to augment training data. We test the effect of this data augmentation on two critical NLP tasks: machine translation and sentiment analysis. The proposed variation generation framework augments the training data with new orthographic variants which are relevant for the test set but did not occur in the training set originally. Our results demonstrate the positive effect of augmenting the training data with a combination of real texts from other corpora as well as synthesized orthographic variation, resulting in performance improvements of 2.1 points in sentiment analysis and 1.4 BLEU points in translation to English.

📄 PDF Abstract BibTeX arXiv:2404.18264

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMachine TranslationSentiment AnalysisTranslation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Low-Resource Cross-Lingual Adaptive Training for Nigerian Pidgin

2023-07-01 · Pin-Jie Lin, Muhammed Saeed, Ernie Chang, Merel Scholman

Developing effective spoken language processing systems for low-resource languages poses several challenges due to the lack of parallel data and limited resources for fine-tuning models. In this work, we target on improv…

text-classificationText ClassificationTranslation

Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus

2024-10-03 · Craig Messner, Tom Lippincott

We present a dataset of 19th century American literary orthovariant tokens with a novel layer of human-annotated dialect group tags designed to serve as the basis for computational experiments exploring literarily meanin…

Language ModelingLanguage Modelling

Semantic Enrichment of Nigerian Pidgin English for Contextual Sentiment Classification

2020-03-27 · Wuraola Fisayo Oyewusi, Olubayo Adekanmbi, Olalekan Akinsande

Nigerian English adaptation, Pidgin, has evolved over the years through multi-language code switching, code mixing and linguistic adaptation. While Pidgin preserves many of the words in the normal English language corpus…

ClassificationGeneral ClassificationSentiment AnalysisSentiment Classification

The Functionalizer: Lossless Functional Decomposition for Subword Tokenization

2026-09-18 · Connor Makowski, Willem Guter hf

Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and Héllo) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation thr…

A Survey of Orthographic Information in Machine Translation

2020-08-04 · Bharathi Raja Chakravarthi, Priya Rani, Mihael Arcan, John P. McCrae

Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor la…

Bilingual Lexicon InductionMachine TranslationSurveyTranslation