paper-with-me

Papers

MaLLaM -- Malaysia Large Language Model

2024-01-26 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

Addressing the gap in Large Language Model pretrained from scratch with Malaysian context, We trained models with 1.1 billion, 3 billion, and 5 billion parameters on a substantial 349GB dataset, equivalent to 90 billion tokens based on our pretrained Byte Pair Encoding (BPE) tokenizer for a single epoch. MaLLaM contributes to enhanced natural language understanding and generation tasks in the Malay language. Although trained on a smaller dataset of 90 billion tokens, our instruction-tuned MaLLaM models perform competitively. When compared to ChatGPT3.5 and Malaysian Mistral, MaLLaM's instruction-tuned models demonstrate notable proficiency, underscoring the effectiveness of our approach in capturing and understanding the nuances of the Malaysian language. MaLLaM models mark a significant contribution to the field, providing comprehensive language representations grounded in Malaysian context. This endeavor aims to pave the way for enhanced natural language understanding and generation tasks specific to the linguistic nuances present in Malaysia. We discuss the training methodology, dataset composition, and the potential impact of MaLLaM in advancing the capabilities of large language models within the context of the Malay language. All models released at https://huggingface.co/collections/mesolitica/mallam-6577b59d1e0b436ae75f930f

📄 PDF Abstract BibTeX arXiv:2401.14680

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelmodelNatural Language Understanding

Similar Papers 제목 키워드 기반

Adapting Safe-for-Work Classifier for Malaysian Language Text: Enhancing Alignment in LLM-Ops Framework

2024-07-30 · Aisyah Razak, Ariff Nazhan, Kamarul Adha, Wan Adzhar Faiq Adzlan 외

As large language models (LLMs) become increasingly integrated into operational workflows (LLM-Ops), there is a pressing need for effective guardrails to ensure safe and aligned interactions, including the ability to det…

Large Malaysian Language Model Based on Mistral for Enhanced Local Language Understanding

2024-01-24 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

In this paper, we present significant advancements in the pretraining of Mistral 7B, a large-scale language model, using a dataset of 32.6 GB, equivalent to 1.1 billion tokens. We explore the impact of extending the cont…

BenchmarkingLanguage ModelingLanguage Modelling

Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations

2024-02-05 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

In this work, we present a comprehensive exploration of finetuning Malaysian language models, specifically Llama2 and Mistral, on embedding tasks involving negative and positive pairs. We release two distinct models tail…

RAGRetrievalRetrieval-augmented GenerationSemantic Similarity+1

Bridging the Gap: Transfer Learning from English PLMs to Malaysian English

2024-07-01 · Mohan Raj Chanthran, Lay-Ki Soon, Huey Fang Ong, Bhawani Selvaretnam

Malaysian English is a low resource creole language, where it carries the elements of Malay, Chinese, and Tamil languages, in addition to Standard English. Named Entity Recognition (NER) models underperform when capturin…

Language Modellingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Malaysian English News Decoded: A Linguistic Resource for Named Entity and Relation Extraction

2024-02-22 · Mohan Raj Chanthran, Lay-Ki Soon, Huey Fang Ong, Bhawani Selvaretnam

Standard English and Malaysian English exhibit notable differences, posing challenges for natural language processing (NLP) tasks on Malaysian English. Unfortunately, most of the existing datasets are mainly based on sta…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+3