Megadose AI progress, ranked and analyzed.

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

· HF Daily Papers ·
The paper says a Qwen2-1.5B-level pretraining run can be brought under $5,090 in fitted cost, using consumer RTX 5090s.

The authors trained Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision. Their best run cost under $6.9K and approached Qwen2.5-1.5B performance under their evaluation setup. A fitted “Puro Cost Scaling Law” suggests about $4.4K is enough to reach Qwen2-1.5B performance. They say the full recipe, data, code, and weights are released under Apache 2.0. HF Daily Papers' note

score 6

Categories: Model Releases, OSS & Tools