Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck
Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing - grounded in a unified framework that decomposes
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming - Nature →
- LinkedLinked via arxiv author · 85%Davide Romano →
“Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck”
- LinkedLinked via arxiv author · 85%Kanak Raj →
“Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck”
- LinkedLinked via arxiv author · 85%Jerrod Parker →
“Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck”
- LinkedLinked via arxiv author · 85%Daniele Giofrè →
“Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck”
