Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models
Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-commit router captures; however, that guarantee holds only under an idealized oracle equipped with correctness labels and an unconstrained budget, neither of which a deployed system has. To the best of our knowledge, no previous work treats resampling the committed model and rerouting to an alternative model as competing
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 50%PacificAI/langtest →
- PossiblePossibly related (embedding) · 49%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 47%I developed a 270 million parameter language model entirely from scratch as an independent research project →
- PossiblePossibly related (embedding) · 47%edwardcapriolo/deliverance →
- LinkedLinked via arxiv author · 85%Teng-Ruei Chen →
“Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models”
