The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language
Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languag
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 53%Large Language Models for Small Language Communities - Times Kuwait →
- PossiblePossibly related (embedding) · 50%The shrinking landscape of linguistic diversity in the age of large language models - Nature →
- PossiblePossibly related (embedding) · 48%Large language models for small language communities are essential - The Business Times →
- LinkedLinked via arxiv author · 85%MohammadHossein Mortazavi →
“The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language”
- LinkedLinked via arxiv author · 85%Mostafa Salehi →
“The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language”
- LinkedLinked via arxiv author · 85%Hadi Veisi →
“The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language”
