Megadose AI progress, ranked and analyzed.

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

· HF Daily Papers ·
The paper reports a 2.93x training-system speedup for full-parameter post-training of DeepSeek-V4 models on Ascend NPUs.

The authors describe a stack of optimizations across parallelism, communication scheduling, and kernel execution for Ascend SuperPOD. They say the system reaches 34.22% MFU while keeping training stable. On top of that, they build a CPT/SFT pipeline for Operations Research tasks using DeepSeek-V4-Flash and solver-verified synthetic data. The resulting specialized model posts a 71.81% zero-shot Pass@1 average in their evaluation, ahead of GPT-5.4-Mini and the base Flash model. HF Daily Papers' note

score 5

Categories: Research