Scalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%llmsresearch/llm-flashcards →
- PossiblePossibly related (embedding) · 50%What exactly does word2vec learn? →
- PossiblePossibly related (embedding) · 49%Does intelligence ‘emerge’ in large language models? - Santa Fe Institute →
- LinkedLinked via arxiv author · 85%Yiming Zhang →
“Scalable Visual Pretraining for Language Intelligence”
- LinkedLinked via arxiv author · 85%Zhonghan Zhao →
“Scalable Visual Pretraining for Language Intelligence”
- LinkedLinked via arxiv author · 85%Wenwei Zhang →
“Scalable Visual Pretraining for Language Intelligence”
- LinkedLinked via arxiv author · 85%Haiteng Zhao →
“Scalable Visual Pretraining for Language Intelligence”
- LinkedLinked via arxiv author · 85%Tianyang Lin →
“Scalable Visual Pretraining for Language Intelligence”
