paper-with-me

홈 › Papers

Aioli: A Unified Optimization Framework for Language Model Data Mixing

2024-11-08 · Mayee F. Chen, Michael Y. Hu, Nicholas Lourie, Kyunghyun Cho, Christopher Ré

Language model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to efficiently learn mixture proportions, ranging from fitting regression models over training runs to dynamically updating proportions throughout training. Surprisingly, we find that no existing method consistently outperforms a simple stratified sampling baseline in terms of average test perplexity. To understand this inconsistency, we unify existing methods into a standard framework, showing they are equivalent to solving a common optimization problem: minimize average loss subject to a method-specific mixing law -- an implicit assumption on the relationship between loss and mixture proportions. This framework suggests that measuring the fidelity of a method's mixing law can offer insights into its performance. Empirically, we find that existing methods set their mixing law parameters inaccurately, resulting in the inconsistent mixing performance we observe. Using this insight, we derive a new online method named Aioli, which directly estimates the mixing law parameters throughout training and uses them to dynamically adjust proportions. Aioli outperforms stratified sampling on 6 out of 6 datasets by an average of 0.27 test perplexity points, whereas existing methods fail to consistently beat stratified sampling, doing up to 6.9 points worse. Moreover, in a practical setting where proportions are learned on shorter runs due to computational constraints, Aioli can dynamically adjust these proportions over the full training run, consistently improving performance over existing methods by up to 12.012 test perplexity points.

📄 PDF Abstract BibTeX arXiv:2411.05735

Code (1)

hazyresearch/aioli 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingMath

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Bridging Large Language Models and Optimization: A Unified Framework for Text-attributed Combinatorial Optimization

2024-08-22 · Xia Jiang, Yaoxin Wu, YuAn Wang, Yingqian Zhang

To advance capabilities of large language models (LLMs) in solving combinatorial optimization problems (COPs), this paper presents the Language-based Neural COP Solver (LNCS), a novel framework that is unified for the en…

Combinatorial OptimizationDecoderLanguage ModellingLarge Language Model

UniAPO: Unified Multimodal Automated Prompt Optimization

2025-08-25 · Qipeng Zhu, Yanzhe Chen, Huasong Zhong, Yan Li 외 arxiv

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in te…

PhaseEvo: Towards Unified In-Context Prompt Optimization for Large Language Models

2024-02-17 · Wendi Cui, Jiaxin Zhang, Zhuohang Li, Hao Sun 외

Crafting an ideal prompt for Large Language Models (LLMs) is a challenging task that demands significant resources and expert human input. Existing work treats the optimization of prompt instruction and in-context learni…

Computational EfficiencyIn-Context Learning

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

2025-05-23 · Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang 외

This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and u…

Image Generationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

promptolution: A Unified, Modular Framework for Prompt Optimization

2025-12-02 · Tom Zehle, Timo Heiß, Moritz Schlager, Matthias Aßenmacher 외 arxiv

Prompt optimization has become crucial for enhancing the performance of large language models (LLMs) across a broad range of tasks. Although many research papers demonstrate its effectiveness, practical adoption is hinde…