Visual General Intelligence: A White Paper
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we dis
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Transformers in Deep Learning: How Self-Attention Changed Modern AI - Snowflake →
- PossiblePossibly related (embedding) · 54%Industry Insights: A Guide to Vision Language Models: Emerging Trends and Applications - A3 Association for Advancing Automation →
- PossiblePossibly related (embedding) · 52%This AI actually learns how your eyes read—what happens next is stunning - Futura, le média qui explore le monde →
- PossiblePossibly related (embedding) · 51%How Iowa State University’s Translational AI Center is feeding the future with generative AI and computer vision on AWS - Amazon Web Services (AWS) →
- PossiblePossibly related (embedding) · 51%Does intelligence ‘emerge’ in large language models? - Santa Fe Institute →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
