Megadose AI progress, ranked and analyzed.

Long-Context Fine-Tuning with Limited VRAM

· ArXiv · AI/CL/LG ·
The paper claims 16K-token fine-tuning on a 16 GB GPU by moving old attention state out of VRAM.

The method combines Hierarchical Global Attention, segment-wise backpropagation, and tiered KV storage. On Qwen3-8B with 4-bit QLoRA, dense training failed at 4,096 tokens on a Quadro RTX 5000, while HGA reached 16,384 tokens at 15.28 GB peak VRAM. The same adapter evaluated through 131,072 tokens on that card, with RAM and NVMe becoming the practical constraint. At 2K training length, HGA and dense adapters scored nearly the same under dense-attention readout: 2.7405 versus 2.7383 nat. ArXiv · AI/CL/LG's note

score 5

Categories: Research