paper-with-me

홈 › Papers

Efficient Reasoning at Fixed Test-Time Cost via Length-Aware Attention Priors and Gain-Aware Training

2026-03-10 · Rian Atri arxiv

We study efficient reasoning under tight compute. We ask how to make structured, correct decisions without increasing test time cost. We add two training only components to small and medium Transformers that also transfer to broader differentiable optimizers. First, a length aware attention prior built via fuzzy regime position alignment, RPA, yields a normalized pre softmax bias that guides attention like a structured regularizer while adding no new inference parameters. Second, a minimal gain aware controller, Guardian, nudges attention sharpness only when validation improvements warrant it, following a two timescale policy gradient view of nonconvex optimization. It is disabled at inference. A KL perspective shows softmax of z plus log pi as MAP with KL regularization, grounding the prior in a principled objective. Under strict compute parity on WikiText 2, we reduce validation cross entropy while matching baseline latency and memory. At inference, we add a precomputed, cached prior B of T as a single additive bias per head. The controller does not run. In practice, this incurs negligible overhead, a cached bias add per head, with no measurable p50 latency shift. Our results suggest that length aware priors and late phase gain control preserve scarce improvements, especially in long span, noisy logit regimes, while keeping test time costs effectively unchanged.

📄 PDF Abstract BibTeX arXiv:2603.09253

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

2025-03-06 · Pranjal Aggarwal, Sean Welleck

Reasoning language models have shown an uncanny ability to improve performance at test-time by ``thinking longer''-that is, by generating longer chain-of-thought sequences and hence using more compute. However, the lengt…

R2-Router: A New Paradigm for LLM Routing with Reasoning

2026-02-02 · Jiaqi Xue, Qian Lou, Jiarong Xing, Heng Huang arxiv

As LLMs proliferate with diverse capabilities and costs, LLM routing has emerged by learning to predict each LLM's quality and cost for a given query, then selecting the one with high quality and low cost. However, exist…

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

2025-05-23 · Jinyan Su, Claire Cardie

Large language models (LLMs) have demonstrated strong reasoning abilities in mathematical tasks, often enhanced through reinforcement learning (RL). However, RL-trained models frequently produce unnecessarily long reason…

Reinforcement Learning (RL)

Reducing Reasoning Costs: The Path of Optimization for Chain of Thought via Sparse Attention Mechanism

2024-11-14 · Libo Wang

In order to address the chain of thought in the large language model inference cost surge, this research proposes to use a sparse attention mechanism that only focuses on a few relevant tokens. The researcher constructed…

Language ModelingLanguage ModellingLarge Language Model

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

2025-06-05 · Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile 외

Large reasoning models (LRMs) achieve higher performance on challenging reasoning tasks by generating more tokens at inference time, but this verbosity often wastes computation on easy problems. Existing solutions, inclu…