A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis of the field remains absent. This study presents a large-scale bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from six collections. We employ BERTopic for topic modeling, regression analysis to identify citation predictors, social network analysis for co-authorship structures, and geographic mapping. O
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 87%MaartenGr/BERTopic →
“Fuzzy title match (0.94): “A Comprehensive Analysis of Arabic Natural Language Processi” ≈ “MaartenGr/BERTopic””
- FuzzySimilar title/name (fuzzy) · 59%google-research/google-research →
“Fuzzy title match (0.73): “A Comprehensive Analysis of Arabic Natural Language Processi” ≈ “google-research/google-research””
- LinkedLinked via arxiv author · 85%Mullosharaf K. Arabov →
“A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A”
