paper-with-me

홈 › Papers

Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets

2024-02-12 · Israel Abebe Azime, Atnafu Lambebo Tonja, Tadesse Destaw Belay, Mitiku Yohannes Fuge, Aman Kassahun Wassie, Eyasu Shiferaw Jada, Yonas Chanie, Walelign Tewabe Sewunetie, Seid Muhie Yimam

Large language models (LLMs) have received a lot of attention in natural language processing (NLP) research because of their exceptional performance in understanding and generating human languages. However, low-resource languages are left behind due to the unavailability of resources. In this work, we focus on enhancing the LLaMA-2-Amharic model by integrating task-specific and generative datasets to improve language model performance for Amharic. We compile an Amharic instruction fine-tuning dataset and fine-tuned LLaMA-2-Amharic model. The fine-tuned model shows promising results in different NLP tasks. We open-source our dataset creation pipeline, instruction datasets, trained models, and evaluation outputs to promote language-specific studies on these models.

📄 PDF Abstract BibTeX arXiv:2402.08015

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages

2024-03-11 · Michael Andersland

Large Language Models (LLMs) like GPT-4 and LLaMA have shown incredible proficiency at natural language processing tasks and have even begun to excel at tasks across other modalities such as vision and audio. Despite the…

BenchmarkingData Augmentation

Semantically Corrected Amharic Automatic Speech Recognition

2024-04-20 · Samuael Adnew, Paul Pu Liang

Automatic Speech Recognition (ASR) can play a crucial role in enhancing the accessibility of spoken languages worldwide. In this paper, we build a set of ASR tools for Amharic, a language spoken by more than 50 million p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSentence+2

The Effect of Normalization for Bi-directional Amharic-English Neural Machine Translation

2022-10-27 · Tadesse Destaw Belay, Atnafu Lambebo Tonja, Olga Kolesnikova, Seid Muhie Yimam 외

Machine translation (MT) is one of the main tasks in natural language processing whose objective is to translate texts automatically from one natural language to another. Nowadays, using deep neural networks for MT tasks…

Machine TranslationSentenceTranslation

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

2023-07-20 · Ephrem Afele Retta, Richard Sutcliffe, Jabar Mahmood, Michael Abebe Berwo 외

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is trained on a pre-existing dataset for that same language. However, where training data for a language does not exist, data fro…

Cross-corpusEmotion RecognitionSpeech Emotion Recognition

Improving Amharic Handwritten Word Recognition Using Auxiliary Task

2022-02-25 · Mesay Samuel Gondere, Lars Schmidt-Thieme, Durga Prasad Sharma, Abiot Sinamo Boltena

Amharic is one of the official languages of the Federal Democratic Republic of Ethiopia. It is one of the languages that use an Ethiopic script which is derived from Gee'z, ancient and currently a liturgical language. Am…

Handwritten Text RecognitionOptical Character RecognitionOptical Character Recognition (OCR)