Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \$1.5M, and reproducing SmolLM3-3B needs over \$700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collec
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Mac Studio M5 Max Cost Analysis →
- PossiblePossibly related (embedding) · 51%Any word on Qwen 3.7 9B? (Also looking for 9B-class alternatives to Qwen 3.5) →
- PossiblePossibly related (embedding) · 49%Don't want to be this guy, but I need Qwen 3.8 35B A3B →
- LinkedLinked via arxiv author · 85%Kairong Luo →
“Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090”
- LinkedLinked via arxiv author · 85%Jiarui Cui →
“Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090”
- LinkedLinked via arxiv author · 85%Yaorui Yin →
“Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090”
- LinkedLinked via arxiv author · 85%Shengqi Chen →
“Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090”
- LinkedLinked via arxiv author · 85%Yiming Yang →
“Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090”
