paper-with-me

홈 › Papers

Lost in Space: Optimizing Tokens for Grammar-Constrained Decoding

2025-02-20 · Sil Hamilton, David Mimno

General-purpose language models are trained to produce varied natural language outputs, but for some tasks like annotation or classification we need more specific output formats. LLM systems increasingly support structured output, sampling tokens according to a grammar, which enforces a format but which can also reduce performance. We ask whether there are systematic differences between grammars that appear semantically similar to humans. To answer this question, we test four popular model families with five token formats on four NLP benchmarks. All models perform most accurately when instructed to classify with real numbers. Performance also improves by 5%-10% when models are instructed to return tokens incorporating leading whitespace, which we find can help models avoid structural deficiencies in subword token representations. Format-based differences are largest for smaller models that are often used for local laptop-scale inference. We present best practices for researchers using language models as zero-shot classifiers with structured output.

📄 PDF Abstract BibTeX arXiv:2502.14969

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accelerating Constrained Decoding with Token Space Compression

2026-05-28 · Michael Sullivan, Alexander Koller arxiv

To guarantee that an LLM's outputs conform to a specified structure, context-free grammar (CFG) decoding engines force the selection of next tokens that produce strings that conform to a given CFG. While current CFG-cons…

Flexible and Efficient Grammar-Constrained Decoding

2025-02-07 · Kanghee Park, Timothy Zhou, Loris D'Antoni

Large Language Models (LLMs) are often asked to generate structured outputs that obey precise syntactic rules, such as code snippets or formatted data. Grammar-constrained decoding (GCD) can guarantee that LLM outputs ma…

XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

2024-11-22 · Yixin Dong, Charlie F. Ruan, Yaxing Cai, Ruihang Lai 외

The applications of LLM Agents are becoming increasingly complex and diverse, leading to a high demand for structured outputs that can be parsed into code, structured function calls, and embodied agent commands. These de…

GPU

Attention Meets Reachability: Structural Equivalence and Efficiency in Grammar-Constrained LLM Decoding

2026-03-04 · Faruk Alpay, Bilge Senturk arxiv

We study grammar-constrained decoding (GCD) as a coupling between an autoregressive next-token distribution and a reachability oracle over a pushdown system compiled from a context-free grammar (CFG). We prove an oracle …

Grammar-Aligned Decoding

2024-05-31 · Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova 외

Large Language Models (LLMs) struggle with reliably generating highly structured outputs, such as program code, mathematical formulas, or well-formed markup. Constrained decoding approaches mitigate this problem by greed…

Code Generation