paper-with-me

Papers

Large Malaysian Language Model Based on Mistral for Enhanced Local Language Understanding

2024-01-24 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

In this paper, we present significant advancements in the pretraining of Mistral 7B, a large-scale language model, using a dataset of 32.6 GB, equivalent to 1.1 billion tokens. We explore the impact of extending the context length, releasing models with context lengths of 4096 and 32768 tokens, and further refining performance with a specialized 16384 context length instruction-tuned model, we called it Malaysian Mistral. Our experiments demonstrate the efficacy of continue pretraining and the influence of extended context lengths on Mistral 7B's language understanding capabilities. Additionally, we release a model specifically tuned with a 16384 context length instruction, showcasing its potential for capturing nuanced language intricacies. Furthermore, our research contributes to the benchmarking of Malaysian Mistral against prominent language models, including ChatGPT3.5 and Claude 2. We present compelling results indicating Malaysian Mistral's superior performance on Tatabahasa (Malay grammar) test set, particularly when fine-tuned with instructions. All models released at https://huggingface.co/collections/mesolitica/malaysian-mistral-7b-6528f2ec825f4bba46c1700c

📄 PDF Abstract BibTeX arXiv:2401.13565

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

MaLLaM -- Malaysia Large Language Model

2024-01-26 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

Addressing the gap in Large Language Model pretrained from scratch with Malaysian context, We trained models with 1.1 billion, 3 billion, and 5 billion parameters on a substantial 349GB dataset, equivalent to 90 billion …

Language ModelingLanguage ModellingLarge Language Modelmodel+1

Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations

2024-02-05 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

In this work, we present a comprehensive exploration of finetuning Malaysian language models, specifically Llama2 and Mistral, on embedding tasks involving negative and positive pairs. We release two distinct models tail…

RAGRetrievalRetrieval-augmented GenerationSemantic Similarity+1

MMMModal -- Multi-Images Multi-Audio Multi-turn Multi-Modal

2024-02-17 · Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

Our contribution introduces a groundbreaking multimodal large language model designed to comprehend multi-images, multi-audio, and multi-images-multi-audio within a single multiturn session. Leveraging state-of-the-art m…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

Adapting Safe-for-Work Classifier for Malaysian Language Text: Enhancing Alignment in LLM-Ops Framework

2024-07-30 · Aisyah Razak, Ariff Nazhan, Kamarul Adha, Wan Adzhar Faiq Adzlan 외

As large language models (LLMs) become increasingly integrated into operational workflows (LLM-Ops), there is a pressing need for effective guardrails to ensure safe and aligned interactions, including the ability to det…

Bridging the Gap: Transfer Learning from English PLMs to Malaysian English

2024-07-01 · Mohan Raj Chanthran, Lay-Ki Soon, Huey Fang Ong, Bhawani Selvaretnam

Malaysian English is a low resource creole language, where it carries the elements of Malay, Chinese, and Tamil languages, in addition to Standard English. Named Entity Recognition (NER) models underperform when capturin…

Language Modellingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2