paper-with-me

홈 › Papers

JetMoE: Reaching Llama2 Performance with 0.1M Dollars

2024-04-11 · Yikang Shen, Zhen Guo, Tianle Cai, Zengyi Qin

Large Language Models (LLMs) have achieved remarkable results, but their increasing resource demand has become a major obstacle to the development of powerful and accessible super-human intelligence. This report introduces JetMoE-8B, a new LLM trained with less than $0.1 million, using 1.25T tokens from carefully mixed open-source corpora and 30,000 H100 GPU hours. Despite its low cost, the JetMoE-8B demonstrates impressive performance, with JetMoE-8B outperforming the Llama2-7B model and JetMoE-8B-Chat surpassing the Llama2-13B-Chat model. These results suggest that LLM training can be much more cost-effective than generally thought. JetMoE-8B is based on an efficient Sparsely-gated Mixture-of-Experts (SMoE) architecture, composed of attention and feedforward experts. Both layers are sparsely activated, allowing JetMoE-8B to have 8B parameters while only activating 2B for each input token, reducing inference computation by about 70% compared to Llama2-7B. Moreover, JetMoE-8B is highly open and academia-friendly, using only public datasets and training code. All training parameters and data mixtures have been detailed in this report to facilitate future efforts in the development of open foundation models. This transparency aims to encourage collaboration and further advancements in the field of accessible and efficient LLMs. The model weights are publicly available at https://github.com/myshell-ai/JetMoE.

📄 PDF Abstract BibTeX arXiv:2404.07413

Code (5)

myshell-ai/jetmoe 공식 구현 pytorch
MS-P3/code3/tree/main/jetmoe mindspore
MS-P3/code5/tree/main/jetmoe mindspore
MindSpore-scientific-2/code-14/tree/main/jetmoe mindspore
pwc-1/Paper-9/tree/main/jetmoe mindspore

Tasks

GPUMixture-of-Experts

Similar Papers 제목 키워드 기반

Detection Made Easy: Potentials of Large Language Models for Solidity Vulnerabilities

2024-09-15 · Md Tauseef Alam, Raju Halder, Abyayananda Maiti

The large-scale deployment of Solidity smart contracts on the Ethereum mainnet has increasingly attracted financially-motivated attackers in recent years. A few now-infamous attacks in Ethereum's history includes DAO att…

Multi-class ClassificationVulnerability Detection

Binary Code Summarization: Benchmarking ChatGPT/GPT-4 and Other Large Language Models

2023-12-15 · Xin Jin, Jonathan Larson, Weiwei Yang, Zhiqiang Lin

Binary code summarization, while invaluable for understanding code semantics, is challenging due to its labor-intensive nature. This study delves into the potential of large language models (LLMs) for binary code compreh…

BenchmarkingCode SummarizationGPUSemantic Similarity+1

MotionLLaMA: A Unified Framework for Motion Synthesis and Comprehension

2024-11-26 · Zeyu Ling, Bo Han, Shiyang Li, Hongdeng Shen 외

This paper introduces MotionLLaMA, a unified framework for motion synthesis and comprehension, along with a novel full-body motion tokenizer called the HoMi Tokenizer. MotionLLaMA is developed based on three core princip…

Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1

Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements

2024-10-22 · Isamu Isozaki, Manil Shrestha, Rick Console, Edward Kim

Hacking poses a significant threat to cybersecurity, inflicting billions of dollars in damages annually. To mitigate these risks, ethical hacking, or penetration testing, is employed to identify vulnerabilities in system…

Llama 3 Meets MoE: Efficient Upcycling

2024-12-13 · Aditya Vavre, Ethan He, Dennis Liu, Zijie Yan 외

Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternative, increasing capacity without a propor…

Mixture-of-ExpertsMMLUMulti-task Language Understanding