paper-with-me

홈 › Papers

Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

2026-08-26 · Ty Chermsirivatana, John MacCormick arxiv

One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the constrained case of this "model size vs. inference compute" trade-off, in which the model outputs are constrained by a strict grammar at inference time. Our results demonstrate that the constrained trade-off behaves differently from the unconstrained trade-off. We investigate the task of converting a prose query into an equivalent SQL query (text-to-SQL). Performance is evaluated on the Spider text-to-SQL benchmark, using the Qwen2.5-Instruct model family ranging in size from 0.5B to 7B parameters, all at 4-bit precision. We experiment with two approaches to varying inference compute: (i) beam search with a variable number of beams; and (ii) sample+vote, i.e., sampling several constrained outputs and then voting on their execution results, where the number of samples is varied. On the 1034-example development set, we find that: (a) both beam search and sample+vote improve accuracy, especially on smaller model sizes; (b) the "model size vs.\ inference compute" trade-off is not advantageous in this experiment, because moving to a larger model size typically results in higher accuracy than increasing inference compute on the same model size; (c) beam search outperforms sample+vote at a matched inference budget. This latter result is of particular interest since it contrasts with the findings of the unconstrained trade-off.

📄 PDF Abstract BibTeX arXiv:2608.25761

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BEAR: Towards Beam-Search-Aware Optimization for Recommendation with Large Language Models

2026-01-30 · Weiqin Yang, Bohao Wang, Zhenxiang Xu, Jiawei Chen 외 arxiv

Recent years have seen a rapid surge in research leveraging Large Language Models (LLMs) for recommendation. These methods typically employ supervised fine-tuning (SFT) to adapt LLMs to recommendation scenarios, and util…

Pushing the Limits of Beam Search Decoding for Transducer-based ASR models

2025-05-30 · Lilit Grigoryan, Vladimir Bataev, Andrei Andrusenko, Hainan Xu 외

Transducer models have emerged as a promising choice for end-to-end ASR systems, offering a balanced trade-off between recognition accuracy, streaming capabilities, and inference speed in greedy decoding. However, beam s…

GPU

Self-Evaluation Guided Beam Search for Reasoning

2023-05-01 · NeurIPS 2023 11 · Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao 외

Breaking down a problem into intermediate steps has demonstrated impressive performance in Large Language Model (LLM) reasoning. However, the growth of the reasoning chain introduces uncertainty and error accumulation, m…

Arithmetic ReasoningGSM8KLanguage ModelingLanguage Modelling+2

First the worst: Finding better gender translations during beam search

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Generating machine translations via beam search seeks the most likely output under a model. However, beam search has been shown to amplify demographic biases exhibited by a model. We aim to address this, focusing on gend…

DiversityRerankingSentenceTranslation

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

2026-07-10 · Spiros Baxevanakis, Peng-Jian Yang arxiv

Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choi…