paper-with-me

홈 › Papers

SeaLLMs -- Large Language Models for Southeast Asia

2023-12-01 · Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Zhiqiang Hu, Chenhui Shen, Yew Ken Chia, Xingxuan Li, Jianyu Wang, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen yang, Chaoqun Liu, Hang Zhang, Lidong Bing

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate.

📄 PDF Abstract BibTeX arXiv:2312.00738

Code (2)

damo-nlp-sg/seallms 공식 구현
DAMO-NLP-SG/SeaExam

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia

2025-11-03 · Chaoqun Liu, Mahani Aljunied, Guizhen Chen, Hou Pong Chan 외 arxiv

We introduce SeaLLMs-Audio, the first large audio-language model (LALM) tailored for multiple Southeast Asian (SEA) languages-Indonesian (id), Thai (th), and Vietnamese (vi)-alongside English (en) and Chinese (zh). Train…

Speech-to-Text TranslationSpeech Emotion RecognitionQuestion AnsweringSpeech Recognition

SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

2024-07-29 · Wenxuan Zhang, Hou Pong Chan, Yiran Zhao, Mahani Aljunied 외

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages unde…

DiversityInstruction FollowingMathematical ReasoningWorld Knowledge

SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages

2025-08-09 · Muhammad Dehan Al Kautsar, Aswin Candra, Muhammad Alif Al Hakim, Maxalmina Satria Kahfi 외 arxiv

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations. To address this gap, we introduce SEADia…

SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

2026-06-02 · Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong, Jian Gang Ngui arxiv

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or …

OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data

2026-02-02 · Tan Sang Nguyen, Muhammad Reza Qorib, Hwee Tou Ng arxiv

Large language models (LLMs) have proven to be effective tools for a wide range of natural language processing (NLP) applications. Although many LLMs are multilingual, most remain English-centric and perform poorly on lo…

Continual Pretraining