Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study
OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on English and Chinese document benchmarks, yet their behaviour on Indic scripts is largely uncharacterised. We benchmark ten systems on Devanagari (Hindi): classical EasyOCR; open VLMs (Qwen2.5-VL-3B, Qwen3-VL-8B, olmOCR-7B); specialised OCR-VLMs (DeepSeek-OCR, Unlimited-OCR); and frontier closed models (Gemini 2.5 Flash, Claude Opus 4.7, GPT-5.5, Mistral OCR), across four synthetic degradation conditions and 300 real printed scans. We report fou
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownFind the best open-source OCR models in one place at Papers with Code [P] →
- LinkedLinked via unknownIs Qwen3-VL-2B the only viable VLM for JSON extraction on a "potato"? →
- LinkedLinked via unknownPP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters →
- LinkedLinked via unknownTurboOCR v3 — high-speed document OCR server (C++/CUDA), ~520 img/s on RTX 5090 →
- PossiblePossibly related (embedding) · 47%Ufonik88/invoice-ocr-app →
- PossiblePossibly related (embedding) · 52%th1nhhdk/local_ai_ocr →
- PossiblePossibly related (embedding) · 49%tesseract-ocr/tesseract →
