paperarXivTrust 82 · PrimaryPublished 2mo agoLive · 3mo ago
Vision-language pretraining at scale
Joint training recipes that align images and text in one embedding space.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Vision-language pretraining at scale” ≈ “VioletVision-3B””
- LinkedLinked via unknownvlm-starter →
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Vision-language pretraining at scale” ≈ “pytorch/vision””
