Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never committing to
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 50%vlm-starter →
- PossiblePossibly related (embedding) · 50%VioletVision-3B →
- LinkedLinked via arxiv author · 85%Spiros Baxevanakis →
“Test-Time Scaling for Small VLMs on Multilingual Visual MCQ”
- LinkedLinked via arxiv author · 85%Peng-Jian Yang →
“Test-Time Scaling for Small VLMs on Multilingual Visual MCQ”
