Read original ↗
paperarXivTrust 82 · PrimaryPublished 2mo agoLive · 3mo ago

Vision-language pretraining at scale

Joint training recipes that align images and text in one embedding space.

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Vision-language pretraining at scale” ≈ “VioletVision-3B”

  • LinkedLinked via unknownvlm-starter
  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Vision-language pretraining at scale” ≈ “pytorch/vision”

Has model

Implements (incoming)

Related across the graph

Topics