paper-with-me

Papers

Very Low Resource Sentence Alignment: Luhya and Swahili

2022-10-01 · loresmt (COLING) 2022 10 · Everlyn Chimoto, Bruce Bassett

Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LASER and LaBSE in extracting bitext for two related low-resource African languages: Luhya and Swahili. For this work, we created a new parallel set of nearly 8000 Luhya-English sentences which allows a new zero-shot test of LASER and LaBSE. We find that LaBSE significantly outperforms LASER on both languages. Both LASER and LaBSE however perform poorly at zero-shot alignment on Luhya, achieving just 1.5% and 22.0% successful alignments respectively (P@1 score). We fine-tune the embeddings on a small set of parallel Luhya sentences and show significant gains, improving the LaBSE alignment accuracy to 53.3%. Further, restricting the dataset to sentence embedding pairs with cosine similarity above 0.7 yielded alignments with over 85% accuracy.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings

Similar Papers 제목 키워드 기반

Very Low Resource Sentence Alignment: Luhya and Swahili

2022-10-31 · Everlyn Asiko Chimoto, Bruce A. Bassett

Language-agnostic sentence embeddings generated by pre-trained models such as LASER and LaBSE are attractive options for mining large datasets to produce parallel corpora for low-resource machine translation. We test LAS…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1

Kencorpus: A Kenyan Language Corpus of Swahili, Dholuo and Luhya for Natural Language Processing Tasks

2022-08-25 · Barack Wanjawa, Lilian Wanzare, Florence Indede, Owen McOnyango 외

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has bee…

Machine TranslationPart-Of-Speech TaggingPOSQuestion Answering+2

COMET-QE and Active Learning for Low-Resource Machine Translation

2022-10-27 · Everlyn Asiko Chimoto, Bruce A. Bassett

Active learning aims to deliver maximum benefit when resources are scarce. We use COMET-QE, a reference-free evaluation metric, to select sentences for low-resource neural machine translation. Using Swahili, Kinyarwanda …

Active LearningLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translation+2

Cross-language Sentence Selection via Data Augmentation and Rationale Training

2021-06-04 · ACL 2021 5 · Yanda Chen, Chris Kedzie, Suraj Nair, Petra Galuščáková 외

This paper proposes an approach to cross-language sentence selection in a low-resource setting. It uses data augmentation and negative sampling techniques on noisy parallel sentence data to directly learn a cross-lingual…

Data AugmentationMachine TranslationRetrievalSentence+2

State of NLP in Kenya: A Survey

2024-10-13 · Cynthia Jayne Amol, Everlyn Asiko Chimoto, Rose Delilah Gesicho, Antony M. Gitau 외

Kenya, known for its linguistic diversity, faces unique challenges and promising opportunities in advancing Natural Language Processing (NLP) technologies, particularly for its underrepresented indigenous languages. This…

Information RetrievalMachine TranslationSentiment Analysisspeech-recognition+3