paper-with-me

홈 › Papers

Aurora:Activating Chinese chat capability for Mixtral-8x7B sparse Mixture-of-Experts through Instruction-Tuning

2023-12-22 · Rongsheng Wang, Haoming Chen, Ruizhe Zhou, Yaofei Duan, Kunyan Cai, Han Ma, Jiaxi Cui, Jian Li, Patrick Cheong-Iao Pang, Yapeng Wang, Tao Tan

Existing research has demonstrated that refining large language models (LLMs) through the utilization of machine-generated instruction-following data empowers these models to exhibit impressive zero-shot capabilities for novel tasks, without requiring human-authored instructions. In this paper, we systematically investigate, preprocess, and integrate three Chinese instruction-following datasets with the aim of enhancing the Chinese conversational capabilities of Mixtral-8x7B sparse Mixture-of-Experts model. Through instruction fine-tuning on this carefully processed dataset, we successfully construct the Mixtral-8x7B sparse Mixture-of-Experts model named "Aurora." To assess the performance of Aurora, we utilize three widely recognized benchmark tests: C-Eval, MMLU, and CMMLU. Empirical studies validate the effectiveness of instruction fine-tuning applied to Mixtral-8x7B sparse Mixture-of-Experts model. This work is pioneering in the execution of instruction fine-tuning on a sparse expert-mixed model, marking a significant breakthrough in enhancing the capabilities of this model architecture. Our code, data and model are publicly available at https://github.com/WangRongsheng/Aurora

📄 PDF Abstract BibTeX arXiv:2312.14557

Code (1)

WangRongsheng/Aurora 공식 구현 pytorch

Tasks

Instruction FollowingMixture-of-ExpertsMMLU

Similar Papers 제목 키워드 기반

Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral

2024-03-04 · Yiming Cui, Xin Yao

Mixtral, a representative sparse mixture of experts (SMoE) language model, has received significant attention due to its unique model design and superior performance. Based on Mixtral-8x7B-v0.1, in this paper, we propose…

Language ModelingLanguage ModellingMixture-of-Experts

Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling

2024-10-22 · Jialong Li, Shreyansh Tripathi, Lakshay Rastogi, Yiming Lei 외

As machine learning models scale in size and complexity, their computational requirements become a significant barrier. Mixture-of-Experts (MoE) models alleviate this issue by selectively activating relevant experts. Des…

AllGPUMixture-of-ExpertsScheduling

Knowledge Fusion of Chat LLMs: A Preliminary Technical Report

2024-02-25 · Fanqi Wan, ZiYi Yang, Longguang Zhong, Xiaojun Quan 외

Recently, FuseLLM introduced the concept of knowledge fusion to transfer the collective knowledge of multiple structurally varied LLMs into a target LLM through lightweight continual training. In this report, we extend t…

Mixtral of Experts

2024-01-08 · Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch 외

We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts). For e…

Code GenerationCommon Sense ReasoningLanguage ModelingLanguage Modelling+4

Does a Global Perspective Help Prune Sparse MoEs Elegantly?

2026-04-08 · Zeliang Zhang, Nikhil Ghosh, Jiani Liu, Bin Yu 외 arxiv

Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by acti…