paper-with-me

홈 › Papers

Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

2024-08-27 · Simran Kaur, Simon Park, Anirudh Goyal, Sanjeev Arora

We introduce Instruct-SkillMix, an automated approach for creating diverse, high quality SFT data. The Instruct-SkillMix pipeline involves two stages, each leveraging an existing powerful LLM: (1) Skill extraction: uses the LLM to extract core "skills" for instruction-following, either from existing datasets, or by directly prompting the model; (2) Data generation: uses the powerful LLM to generate (instruction, response) data that exhibit a randomly chosen pair of these skills. Here, the use of random skill combinations promotes diversity and difficulty. Vanilla SFT (i.e., no PPO, DPO, or RL methods) on data generated from Instruct-SkillMix leads to strong gains on instruction following benchmarks such as AlpacaEval 2.0, MT-Bench, and WildBench. With just $4$K examples, LLaMA-3-8B-Base achieves 42.76% length-controlled win rate on AlpacaEval 2.0. To our knowledge, this achieves state-of-the-art performance among all models that have only undergone SFT (no RL methods) and competes with proprietary models such as Claude 3 Opus and LLaMA-3.1-405B-Instruct. Ablation studies also suggest plausible reasons for why creating open instruction-tuning datasets via naive crowd-sourcing has proved difficult. Introducing low quality answers ("shirkers") in $20\%$ of Instruct-SkillMix examples causes performance to plummet, sometimes catastrophically. The Instruct-SkillMix pipeline is flexible and is adaptable to other settings.

📄 PDF Abstract BibTeX arXiv:2408.14774

Code (1)

princeton-pli/Instruct-SkillMix 공식 구현 pytorch

Tasks

Instruction Following

Methods 이 논문이 사용한 방법론

DPO 설명 없음
Entropy Regularization 설명 없음
SFT Shrink and Fine-Tune, or SFT, is a type of distillation that avoids explicit distillation by copying parameters to a student student model and then fine-tuning.…
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…

Similar Papers 제목 키워드 기반

EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning

2024-04-25 · CVPR 2024 1 · HongXia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen 외

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural…

Emotion ClassificationEmotion Recognition

ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

2024-12-09 · Jieyu Zhang, Le Xue, Linxin Song, Jun Wang 외

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existing practices rely on powerful but costly …

Graph GenerationScene Graph GenerationVisual Question Answering

Biology Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models

2024-12-26 · Haonan He, Yuchen Ren, Yining Tang, Ziyang Xu 외

Large language models have already demonstrated their formidable capabilities in general domains, ushering in a revolutionary transformation. However, exploring and exploiting the extensive knowledge of these models to c…

Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models

2023-05-24 · Sheng Shen, Le Hou, Yanqi Zhou, Nan Du 외

Sparse Mixture-of-Experts (MoE) is a neural architecture design that can be utilized to add learnable parameters to Large Language Models (LLMs) without increasing inference cost. Instruction tuning is a technique for tr…

Mixture-of-ExpertsZero-shot Generalization

Harnessing the Power of David against Goliath: Exploring Instruction Data Generation without Using Closed-Source Models

2023-08-24 · Yue Wang, Xinrui Wang, Juntao Li, Jinxiong Chang 외

Instruction tuning is instrumental in enabling Large Language Models~(LLMs) to follow user instructions to complete various open-domain tasks. The success of instruction tuning depends on the availability of high-quality…