paper-with-me

Papers

Unsupervised Data Augmentation for Less-Resourced Languages with no Standardized Spelling

2019-09-01 · RANLP 2019 9 · Alice Millour, Kar{\"e}n Fort

Building representative linguistic resources and NLP tools for non-standardized languages is challenging: when spelling is not determined by a norm, multiple written forms can be encountered for a given word, inducing a large proportion of out-of-vocabulary words. To embrace this diversity, we propose a methodology based on crowdsourced alternative spellings we use to extract rules applied to match OOV words with one of their spelling variants. This virtuous process enables the unsupervised augmentation of multi-variant lexicons without expert rule definition. We apply this multilingual methodology on Alsatian, a French regional language and provide an intrinsic evaluation of the correctness of the variants pairs, and an extrinsic evaluation on a downstream task. We show that in a low-resource scenario, 145 inital pairs can lead to the generation of 876 additional variant pairs, and a diminution of OOV words improving the part-of-speech tagging performance by 1 to 4{\%}.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDiversityPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Comparing Approaches to Automatic Summarization in Less-Resourced Languages

2025-12-30 · Chester Palen-Michel, Constantine Lignos arxiv

Automatic text summarization has achieved high performance in high-resourced languages like English, but comparatively less attention has been given to summarization in less-resourced languages. This work compares a vari…

Text SummarizationData Augmentation

Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract

2020-03-30 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya 외

The recent advances in Natural Language Processing have only been a boon for well represented languages, negating research in lesser known global languages. This is in part due to the availability of curated data and res…

Data AugmentationGeneral ClassificationTopic Classification

Weakly-supervised Deep Cognate Detection Framework for Low-Resourced Languages Using Morphological Knowledge of Closely-Related Languages

2023-11-09 · Koustava Goswami, Priya Rani, Theodorus Fransen, John P. McCrae

Exploiting cognates for transfer learning in under-resourced languages is an exciting opportunity for language understanding tasks, including unsupervised machine translation, named entity recognition and information ret…

Information RetrievalMachine Translationnamed-entity-recognitionNamed Entity Recognition+2

A language-independent and fully unsupervised approach to lexicon induction and part-of-speech tagging for closely related languages

2014-05-01 · LREC 2014 5 · Yves Scherrer, Beno{\^\i}t Sagot

In this paper, we describe our generic approach for transferring part-of-speech annotations from a resourced language towards an etymologically closely related non-resourced language, without using any bilingual (i.e., p…

Part-Of-Speech TaggingPOSTranslation

Intent Recognition and Unsupervised Slot Identification for Low Resourced Spoken Dialog Systems

2021-04-03 · Akshat Gupta, Olivia Deng, Akruti Kushwaha, Saloni Mittal 외

Intent Recognition and Slot Identification are crucial components in spoken language understanding (SLU) systems. In this paper, we present a novel approach towards both these tasks in the context of low resourced and un…

Data AugmentationGeneral Classificationintent-classificationIntent Classification+3