Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models
One common trade-off in the use of large language models involves reducing the size of the model while increasing the amount of computation at inference time, for example by using a wider beam search. In this paper, we examine the constrained case of this "model size vs. inference compute" trade-off, in which the model outputs are constrained by a strict grammar at inference time. Our results demonstrate that the constrained trade-off behaves differently from the unconstrained trade-off. We investigate the task of converting a prose query into an equivalent SQL query (text-to-SQL). Performance
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%The Engine Beneath the Oracle: How Large Language Models Actually Work - www.lvivherald.com →
- PossiblePossibly related (embedding) · 52%If you had a 300M parameter model, what would you optimize it for? →
- PossiblePossibly related (embedding) · 50%Large Language Models: Qwen3 Offers AI Models For Deeper Reasoning And Faster Responses - Trend Hunter →
- FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference →
“Fuzzy title match (0.92): “Beam Search, Self-Consistency, and the Limits of Inference-T” ≈ “xorbitsai/inference””
- LinkedLinked via arxiv author · 85%Ty Chermsirivatana →
“Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Lan”
- LinkedLinked via arxiv author · 85%John MacCormick →
“Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Lan”
