Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters →
- PossiblePossibly related (embedding) · 49%Perplexity AI Language Support for Urdu and Arabic Speakers - autogpt.net →
- LinkedLinked via arxiv author · 85%Pouria Mahdi →
“Persian Pixel: A large-scale synthetic OCR dataset for Persian language”
- LinkedLinked via arxiv author · 85%Haq Nawaz Malik →
“Persian Pixel: A large-scale synthetic OCR dataset for Persian language”
