paper-with-me

Papers

Compute Where it Counts: Self Optimizing Language Models

2026-05-11 · Yash Akhauri, Mohamed S. Abdelfattah arxiv

Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attention), typically applying a uniform computation budget to every generated token. In practice, token difficulty varies widely, so static compression can over-compute on easy steps and under-compute on hard ones. We study dynamic budget allocation for autoregressive decoding: learning how much computation to spend per token from within a single model. Self-Optimizing Language Models (SOL) pair a frozen LLM with a lightweight policy network that reads the LLM hidden state and selects a discrete efficiency action at each decode step. Actions can jointly control (i) token-level attention sparsity, (ii) structured activation pruning in the MLP, and (iii) activation quantization bit-width, while leaving the base model weights unchanged. We train the policy with group-relative policy optimization on teacher-forced episodes: the token sequence is fixed, while we sample multiple compute schedules (i.e., "counterfactual" schedules that vary only the efficiency actions for the same token path) and compare their likelihoods under the same supervision. Our reward trades off language-model quality against soft penalties that encourage episode-average budget usage to match a requested target. Across model variants and compute regimes, SOL improves quality at matched budget over static allocation and strong random schedule search, offering a complementary axis for inference-efficiency optimization. SOL discovers a better quality-efficiency pareto-front across all our experiments and improves MMLU accuracy by up to 7.3% over uniform budget allocation strategies.

📄 PDF Abstract BibTeX arXiv:2605.10875

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Rewarding Vision-Language Model via Reasoning Decomposition

2025-08-27 · Zongxia Li, Wenhao Yu, Chengsong Huang, Zhenwen Liang 외 arxiv

Vision-Language Models (VLMs) often suffer from visual hallucinations: generating things that are not consistent with visual inputs and language shortcuts, where they skip the visual part and just rely on text priors. Th…

Reinforcement LearningVisual Reasoning

Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality

2025-08-24 · Daniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato 외 arxiv

Communication locality plays a key role in the performance of collective operations on large HPC systems, especially on oversubscribed networks where groups of nodes are fully connected internally but sparsely linked thr…

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking

2026-05-15 · Vaidehi Bagaria, Nikshep Grampurohit, Pulkit Verma arxiv

Reinforcement learning (RL) allows vision-language-action (VLA) policies to generalize beyond their training distribution by optimizing directly for task success, but post-training is computationally expensive. A natural…

Reinforcement Learning

FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration

2025-12-12 · Dongwon Jung, Peng Shi, Muhao Chen, Yi Zhang arxiv

Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training. However, extending these techniques to multi-agent systems remains challenging: ex…

An Efficient Data Reuse with Tile-Based Adaptive Stationary for Transformer Accelerators

2025-03-25 · Tseng-Jen Li, Tian-Sheuan Chang

Transformer-based models have become the \textit{de facto} backbone across many fields, such as computer vision and natural language processing. However, as these models scale in size, external memory access (EMA) for we…