Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
The paper says a Qwen2-1.5B-level pretraining run can be brought under $5,090 in fitted cost, using consumer RTX 5090s.
The authors trained Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision. Their best run cost under $6.9K and approached Qwen2.5-1.5B performance under their evaluation setup. A fitted “Puro Cost Scaling Law” suggests about $4.4K is enough to reach Qwen2-1.5B performance. They say the full recipe, data, code, and weights are released under Apache 2.0. HF Daily Papers' note
The authors trained Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision. Their best run cost under $6.9K and approached Qwen2.5-1.5B performance under their evaluation setup. A fitted “Puro Cost Scaling Law” suggests about $4.4K is enough to reach Qwen2-1.5B performance. They say the full recipe, data, code, and weights are released under Apache 2.0. HF Daily Papers' note
score 6